One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
guidesMay 13, 2026· 5 min read

How to cache Pydantic AI LLM calls with Crowkis

Add a semantic cache to Pydantic AI so repeated and reworded questions are served for free, no rewrite, self-hosted.

The cheapest token is the one you never spend twice. If you build with Pydantic AI, most of that repetition is invisible in your code but very visible on your bill. A semantic cache in front of your model calls fixes it.

The lowest-friction path is the OpenAI-compatible gateway: point Pydantic AI's base URL at Crowkis and every model call flows through a semantic cache. Repeated and reworded prompts are served from cache with no upstream call; new ones pass through and get cached.

Pydantic AI + Crowkis gateway
# point Pydantic AI at the Crowkis gateway
base_url = "http://127.0.0.1:6380/v1"   # semantic cache in front of your provider
In plain words: You don't restructure your Pydantic AI app. You change where the calls go, and repeats stop costing money.

On repetitive workloads this cuts LLM costs up to 60-70% on repetitive workloads, and every hit comes back with a confidence score so reuse stays safe. Drop it in over RESP, gRPC, REST, or MCP, no rewrite required.