Samuel Agbede / An open notebookSay hello ↗
← Back to notes

Note 5 Oct 2026 15 min read Retrieval that works · part 4 of 4

Using Jev to verify cached LLM answers

I compared three ways of using Jev to check cached answers, then tested it against GPT-4.1 mini. The checks were cheaper and faster, with a trade-off in how many answers could be reused.


Could Jev help reduce LLM costs by checking when a cached answer can be reused?

I wrote about the basics and practical trade-offs in Semantic Caching in Production. This post looks more closely at one part of that system: checking whether a cached answer really answers a new question.

I had been using a lightweight LLM, GPT-4.1 mini, to check answers coming back from a semantic cache. It helped with a recurring problem: two questions can look similar and still need different answers. But that verification has its own cost and latency.

I wondered if Jev could be my new verifier.

So I tried three ways of giving it the task, then compared the most useful approach with GPT-4.1 mini. In this small synthetic test, my Jev setup approved 80% of the valid reuse opportunities and blocked all the unsuitable answers. Its verification checks cost 82% less, with a median response time of 240 ms versus 532 ms.

There was a trade-off. GPT-4.1 mini recognised more answers that could be reused. Jev’s confidence threshold blocked some perfectly usable answers along with the unsuitable ones, which would send more requests to the answering LLM.

Here’s what I tested, where those differences came from, and what I think they mean for an application.

If you want the numbers first, jump to the three-input comparison or Jev versus GPT-4.1 mini.

What Jev does

For those who missed the news, Jev is a model from TypeSafe AI built to make quick structured decisions you can integrate into your application code. TypeSafe calls it a System One model.

You give it some context, called the state, and define the questions you want it to answer. It has three primitives:

  • Choice: chooses an option from a list.
  • Score: returns a score against a rubric.
  • Noul: returns a probability between 0 and 1 for whether a statement is true.

Choice and Score also return a probability distribution and a confidence value. These give your application a way to decide whether to act on the result. TypeSafe’s introduction explains the primitives.

I spent some time learning about its use cases and playing with it. Things like model routing, classification and guardrails made sense, but one use case I found particularly interesting was combining it with semantic caching.

Why a cache hit needs another check

Semantic caching lets you reuse a previous LLM response when a new question is similar enough to one you’ve already answered. In my application, Redis searches for a cached question using the new question’s embedding and returns a candidate answer.

The difficult part is deciding whether that answer actually covers the new request.

Take this example:

Original question: Can I transfer my ticket?

New question: How do I transfer my ticket, and what is the deadline?

The questions are closely related. But whether we can reuse the answer depends on what the original response said.

Cached answerDoes it answer the new question?
Yes, tickets are transferable.No. It leaves out the procedure and deadline.
Yes. Select “Transfer ticket” under “Manage booking” before Friday at 17:00.Yes. It already contains the information the new question asks for.

Both answers fit the original question. Only one fits the new question.

That is the kind of false positive I wanted to guard against: a similarity search finding a related question whose answer cannot be served unchanged. I had experimented with cross-encoders and lightweight LLMs for this, and eventually settled on an LLM verifier after each candidate cache hit.

Where Jev fits in the flow

The flow I wanted to evaluate was straightforward:

  1. A new question arrives and the application searches the semantic cache.
  2. If there is no candidate, the answering LLM generates a fresh response.
  3. If there is a candidate, Jev checks whether the cached answer fully answers the new question.
  4. The application either reuses that answer or calls the answering LLM.

A new question goes to the Redis semantic cache. A candidate answer goes to Jev with the new question. Reuse requires accept and confidence of at least 0.50. A cache miss, rejection, uncertainty or low confidence goes to the answering LLM.

View the diagram at full size.

The diagram shows the intended application flow. The benchmark below measured the Jev verification step; it did not run the cache lookup or generate fallback answers. API and validation failures also require a fresh answer.

The question for Jev was: Can this exact cached answer be served unchanged to fully answer the new question?

I used the Choice primitive with three options:

  • accept: the answer fully covers the new request without changes.
  • reject: the answer is incomplete, mismatched, or assumes something the new request did not say.
  • uncertain: there is not enough information to decide.

After validating the response, my application applied this rule:

reuse_answer = (
    result.choice == "accept"
    and result.confidence >= 0.50
)

An accept decision alone wasn’t enough. Rejection, uncertainty, low confidence or an invalid response all meant falling back to the answering LLM.

Three ways of asking Jev

Before comparing Jev with an LLM, I wanted to understand what information Jev needed. I tried three approaches.

1. Send both questions and the cached answer

The first approach supplied the original cached question, the new question and the cached answer. It asked two separate questions:

  • Do the two questions request equivalent information?
  • Does the cached answer fully answer the new question?

Both judgments had to return accept with confidence of at least 0.50. Jev evaluated both in one request.

This gives the verifier the most context, but the acceptance rule is restrictive. In the ticket example, the questions ask for different information even when the longer cached answer already covers the new request. Requiring question equivalence can reject that useful answer.

2. Send only the two questions

The second approach supplied the original and new questions. Jev checked whether they requested equivalent information and could share an answer.

This is a smaller input, particularly when cached answers are long. But it cannot inspect the answer the user would actually receive. Both rows in the ticket example look identical to this verifier, because it only sees the questions.

3. Send the new question and cached answer

The third approach supplied only the new question and the actual cached answer. Jev checked whether that answer covered the entire new request, including its entities, constraints, negation and time scope.

This directly tests the reuse decision I care about. It also has a limitation: some answers need their original question to make sense. An answer like “Yes, it includes lunch” does not tell you which ticket type “it” refers to. I’ll come back to that.

How I tested the three approaches

I used 60 unique synthetic cases, split evenly between reusable and unsuitable answers. Each approach ran three times on every case, giving 180 calls per approach and 540 Jev calls in total.

Each approach therefore faced 90 valid reuse opportunities and 90 unsuitable-answer trials. Those are repeated trials on 30 cases in each group, rather than 90 independent examples.

The cases covered four situations:

SituationUnique cases
Related questions paired with complete or incomplete answers28
Ordinary paraphrases that could reuse the answer16
Explicit changes to dates, quantities, negation or other scope8
Answers whose meaning depends on the original question8

I fixed the prompts, labels and existing 0.50 threshold before running the comparison. Jev received only the fields allowed by each approach. It did not receive the expected label, the explanation for that label, or the source evidence used to define the synthetic case.

These were tests of candidate answers. I did not run Redis searches or embedding calls to produce the candidates, and I did not execute the fallback generation that a rejected answer would require.

Checking the answer recovered more useful reuses

Here are the results with the confidence rule applied. All costs are in US dollars.

Jev approachValid reuse approvedUnsuitable answers reusedMedian latencyCost for 180 checks
Both questions + answer; both checks required48/90 — 53.3%0/90241 ms$0.005484
Both questions only48/90 — 53.3%0/90235 ms$0.003723
New question + answer72/90 — 80%0/90240 ms$0.003857

Sending the new question and cached answer performed best overall on this mix of cases. It recovered 24 additional valid reuse opportunities compared with either approach that required question equivalence.

The difference came from answers that already covered a broader request. There were 42 reusable-answer trials in that group. The first two approaches rejected all of them; the third recovered 31 of 42.

The pattern changed for ordinary paraphrases. The first two approaches approved all 48 of 48 valid trials in that group, while the third approved 41 of 48.

So I wouldn’t take this as a reason to always remove the original question. Approach 1 changed both the available context and the acceptance rule. I didn’t include a fourth approach that supplied all three fields but made answer applicability the only condition for reuse. That would be a useful next comparison, especially for answers that depend on their original context.

The latency differences between the three approaches were small: less than 6 ms between their measured medians. This run doesn’t establish a meaningful speed winner among them.

The confidence threshold made a real difference

The zero incorrect-reuse result needs some explanation.

For the new-question-plus-answer approach, Jev returned a raw accept on 12 of the 90 unsuitable-answer trials. Every one had confidence below 0.50, so none was reused. One also failed response validation.

New question + answerRaw accept decisionsAnswers actually reused
Valid reuse opportunities84/9072/90
Unsuitable answers12/900/90

The same threshold also blocked 12 raw accepts on valid answers. Along with six valid trials that Jev did not accept in the first place, that left 18 missed reuse opportunities.

That is the trade-off I would expect to investigate when choosing a threshold: how many unsuitable answers get through, and how many usable answers get turned down?

The confidence field is separate from the cache’s embedding similarity and from probabilities.accept. TypeSafe derives confidence from the shape of the returned probability distribution. A threshold of 0.50 is an application decision rule; it does not mean a 50% accuracy guarantee. TypeSafe’s confidence documentation explains the distinction.

I kept the application’s existing threshold for this experiment. I didn’t tune it to produce the zero-incorrect-reuse result, and this test doesn’t establish that it is the best threshold for another workload.

Comparing Jev with GPT-4.1 mini

The next question was whether Jev could do this job more cheaply than the LLM I was already using.

I ran GPT-4.1 mini on the same 60 cases, with three repeats. It received the same new question, cached answer, applicability instructions and decision criteria as Jev’s third approach.

I asked it for a compact structured response containing accept, reject or uncertain. It did not generate an explanation. The setup used strict JSON Schema, temperature 0 and a 64-token output limit.

This gave me a comparison of the same task and inputs. It wasn’t a replay of my application’s full LLM verifier, which also receives the original question and produces a brief reason.

MetricJev with confidence ≥ 0.50GPT-4.1 mini
Valid reuse approved72/90 — 80%90/90 — 100%
Missed valid reuse18/900/90
Unsuitable answers blocked90/9081/90
Unsuitable answers reused0/909/90
Median verification latency240 ms532 ms
p95 verification latency350 ms801 ms
Cost for 180 verification checks$0.003857$0.021720

Jev’s verification cost was 82% lower, and its median verification latency was 55% lower in these runs. GPT-4.1 mini approved every valid reuse opportunity, but also approved nine unsuitable-answer trials.

The decision policies matter here. Jev required both an accept decision and enough confidence. GPT-4.1 mini reused on accept, with rejection, uncertainty or errors requiring fallback. I didn’t ask the LLM to invent a confidence value that would look numerically comparable to Jev’s field.

This compares those two configurations. A different abstention policy for the LLM could change the result.

Where GPT-4.1 mini allowed unsuitable answers

Its nine incorrect reuses came from three distinct cases, each repeated three times.

One involved negation:

New question: Which Maple sessions can I attend without advance booking?

Cached answer: The Maple pottery and welding sessions require advance booking.

The answer names sessions in the category the user wants to exclude. It never identifies the sessions that can be attended without booking. GPT-4.1 mini accepted it in all three repeats; Jev rejected it in all three.

The other two cases involved missing context. For example:

New question: Does the Spruce standard pass include lunch?

Cached answer: Yes, it includes lunch.

The cached answer originally referred to a VIP pass. In the synthetic scenario, the standard pass did not include lunch. But neither verifier received the original question, so that distinction was missing from the input.

GPT-4.1 mini accepted the answer. Jev also returned raw accepts, but its confidence stayed below the threshold and reuse was blocked. The third case had the same problem with an answer about a keynote recording being reused for a private workshop.

These examples make me cautious about throwing away the original question. The application rule happened to block reuse here, but retaining the answer’s original scope would give the verifier information it otherwise cannot recover.

Cheaper verification is only part of the cost

The 82% figure describes the cost of checking candidate answers. It doesn’t tell me how much the entire application would save.

Every missed valid reuse means paying to generate an answer that was already available. My Jev setup missed 18 such opportunities. It would require fallback on 108 of the 180 trials, compared with 81 for GPT-4.1 mini. Nine of the LLM’s apparent avoided fallbacks, however, came from reusing unsuitable answers.

To be honest, I prefer making another generation call to serving an unsuitable answer. But I still want to measure what that preference costs.

For the full application, the calculation needs to include the cache lookup and embedding costs, verification costs, the number and cost of fresh generations, and the consequences of incorrect reuse. End-to-end latency also changes when a verifier rejects a candidate and the user then waits for generation.

This experiment measured only one part of that system. I would need a run through the complete application, using representative requests, to claim overall savings.

Experiment details and limits

I ran these tests on 5 October 2026, using typesafe/jev-1.13-20260917 through OpenRouter and gpt-4.1-mini-2025-04-14 through OpenAI.

The cases were synthetic, some came from earlier evaluations, and repeated calls weren’t independent examples. Zero incorrect reuses here doesn’t establish a zero error rate in production.

The runs happened about an hour apart through different providers, so latency includes network and service conditions. Jev’s cost was provider-reported, including the response that failed validation. GPT-4.1 mini’s was calculated from returned token usage at OpenAI’s published prices.

I tested whether an answer fits a new request. Factual accuracy, freshness and user permissions still need separate checks.

What I want to explore next

The most useful finding for me was how much the question I asked the verifier changed the outcome. Checking whether two questions are equivalent can miss a cached answer that already covers both. Checking the answer directly recovered more of those opportunities in this test.

I want to take this further in a few directions: retain the original question as context without requiring question equivalence, evaluate the confidence threshold on separate data, and measure the complete application’s cost and latency. I’m also curious about which domains Jev struggles with and how it compares with an ML classifier trained specifically for this task.

For now, Jev looks promising as the step between finding a cached answer and deciding to serve it. My setup made that check cheaper and faster, while accepting fewer reuse opportunities. Whether that’s a good trade depends on the requests an application gets and the cost of getting the decision wrong.

I’m exploring this further in a webinar on semantic caching with Jev on Wednesday 7 October 2026 at 2pm BST. I’ll walk through the experiment and the trade-offs involved in deciding when to reuse an answer. Register if you’d like to join.

If you’re using Jev, or checking semantic cache hits another way, I’d love to hear what you’ve found.

Do you have any thoughts, corrections or questions? I'd love to hear from you.

Reply by email samuelagbede@outlook.com