Five TTL policies: engineering the shelf life of truth
Answers age at different speeds, prices in days, math never. A single TTL knob can't express that, so Crowkis ships five policies plus version pinning and webhooks.
Notes from the nest · 980 posts
Engineering notes written by the people building Crowkis. Comparisons, use cases, economics, internals, security, operations, and nothing written just to rank.
Answers age at different speeds, prices in days, math never. A single TTL knob can't express that, so Crowkis ships five policies plus version pinning and webhooks.
Air-gapped networks, FedRAMP postures, and zero phone-home tolerance rule out most AI infrastructure on page one. Crowkis was designed to pass that page.
Building API documentation bots on LangChain.js? Add a semantic cache so the same endpoint questions from every developer stop costing full price.
Building AI search on the OpenAI Python SDK? Add a semantic cache so popular queries hit again and again stop costing full price.
Building voice assistants on DSPy? Add a semantic cache so latency-sensitive, repetitive spoken queries stop costing full price.
Building sales enablement tools on LangGraph? Add a semantic cache so reps asking the same product questions stop costing full price.
Add a semantic cache to the Vercel AI SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
A semantic cache slowly accumulates near-duplicate answers. CDEDUP finds the clusters that mean the same thing and collapses them, reclaiming memory, and Crowkis is honest about its cost.
Exciting infrastructure is a contradiction in terms. Every Crowkis design decision optimizes for the same review: 'it just runs.'
Three sentences, one dashboard number, and a flat price. The rare infrastructure purchase that finance understands faster than engineering does.
Building customer support bots on Spring AI? Add a semantic cache so repeat questions from every customer, all day stop costing full price.
Building ecommerce assistants on the OpenAI Python SDK? Add a semantic cache so the same product and policy questions across shoppers stop costing full price.
Building email drafting tools on DSPy? Add a semantic cache so similar drafts requested over and over stop costing full price.
Building knowledge base assistants on LangGraph? Add a semantic cache so the same lookups across a team all day stop costing full price.
Durable, per-user memory for the Vercel AI SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
'Many eyes' assumes the eyes show up. For your hot path, a signed single binary with zero dependencies is a smaller attack surface than a thousand auditable packages nobody audits.
AWS will happily run an exact-match cache for you at any scale. It will miss your LLM traffic at any scale, too.
Building coding assistants on Spring AI? Add a semantic cache so the same explanations and boilerplate reasoning, dozens of times a day stop costing full price.
Building healthcare Q&A assistants on the OpenAI Python SDK? Add a semantic cache so recurring policy and triage questions stop costing full price.
Building code review bots on DSPy? Add a semantic cache so the same review patterns across pull requests stop costing full price.