Tag: semantic-caching
-
Using Jev to verify cached LLM answers
I compared three ways of using Jev to check cached answers, then tested it against GPT-4.1 mini. The checks were cheaper and faster, with a trade-off in how many answers could be reused.
-
Borrowed confidence is fragile in agentic systems
Production-like evals revealed the retrieval architecture I actually needed and reminded me that confidence in agentic systems has to be earned, not borrowed.
-
Semantic Caching in Production
Repeated user intents can quietly inflate LLM cost and latency. Semantic caching helps, but production use comes with trade-offs.