Selected noteReading preview
5 Oct 2026
Using Jev to verify cached LLM answers
I compared three ways of using Jev to check cached answers, then tested it against GPT-4.1 mini. The checks were cheaper and faster, with a trade-off in how many answers could be reused.
Read in the notebook → 13 Sept 2026
How Does the Dot Product Measure Similarity?
A step-by-step journey from vectors and the law of cosines to why the dot product of normalised vectors measures directional similarity.
Read in the notebook → 28 Aug 2026
How LLMs Turn Text into Embeddings
A first-principles look at how GPT-style language models turn raw text into token IDs, token embeddings, and positional embeddings.
Read in the notebook → 13 Aug 2026
Designing AI Agents from the Outside In
The PEAS framework offers a way to design AI agents from the outside in by defining success, environment, actions, and observations before writing code.
Read in the notebook → 30 Jul 2026
Could agents learn to work better from their own runs?
Most research on agent memory focuses on remembering the user. I built a browser agent that learned from its own runs, and the lessons that transferred surprised me.
Read in the notebook → 10 Jul 2026
Context engineering needs a context engine
Stronger models help, but agent reliability may depend just as much on better context systems around them.
Read in the notebook → 6 Jul 2026
Borrowed confidence is fragile in agentic systems
Production-like evals revealed the retrieval architecture I actually needed and reminded me that confidence in agentic systems has to be earned, not borrowed.
Read in the notebook → 19 Jun 2026
The lethal trifecta in AI agents
When agents can read private data, process untrusted content, and communicate outward, prompt injection becomes a much more serious security problem.
Read in the notebook → 17 Jun 2026
How Plan Caching Reduces LLM Agent Costs
Plan caching reuses planning templates across similar agent tasks, cutting cost and latency without throwing away accuracy.
Read in the notebook → 5 Jun 2026
Stop filling your agent's context window just because you can
Bigger context windows do not remove failure modes. They create new ones when we stop being intentional about what goes into an agent's context.
Read in the notebook → 21 May 2026
I benchmarked 5 embedding models across 4 datasets
I benchmarked five embedding models across four NanoBEIR datasets and found that bigger embeddings did not always produce better retrieval.
Read in the notebook → 19 May 2026
Why reranking matters with cross-encoders
Bi-encoders make retrieval fast, but cross-encoders expose why reranking matters when meaning depends on the query.
Read in the notebook → 11 May 2026
A beginner-friendly guide to the GGUF model format
GGUF made local LLM inference feel practical by packaging model weights, vocabulary, hyperparameters, and architecture metadata into one runnable format.
Read in the notebook → 27 Mar 2026
Semantic Caching in Production
Repeated user intents can quietly inflate LLM cost and latency. Semantic caching helps, but production use comes with trade-offs.
Read in the notebook → Learning Out Loud04