Tag: evaluation
-
Using Jev to verify cached LLM answers
I compared three ways of using Jev to check cached answers, then tested it against GPT-4.1 mini. The checks were cheaper and faster, with a trade-off in how many answers could be reused.
-
Could agents learn to work better from their own runs?
Most research on agent memory focuses on remembering the user. I built a browser agent that learned from its own runs, and the lessons that transferred surprised me.
-
I benchmarked 5 embedding models across 4 datasets
I benchmarked five embedding models across four NanoBEIR datasets and found that bigger embeddings did not always produce better retrieval.