Samuel Agbede / An open notebookSay hello ↗
The technical notebookAll notes

Things I’m
figuring out.

  1. 5 Oct 2026

    Using Jev to verify cached LLM answers

  2. 13 Sept 2026

    How Does the Dot Product Measure Similarity?

  3. 28 Aug 2026

    How LLMs Turn Text into Embeddings

  4. 13 Aug 2026

    Designing AI Agents from the Outside In

  5. 30 Jul 2026

    Could agents learn to work better from their own runs?

  6. 10 Jul 2026

    Context engineering needs a context engine

  7. 6 Jul 2026

    Borrowed confidence is fragile in agentic systems

  8. 19 Jun 2026

    The lethal trifecta in AI agents

  9. 17 Jun 2026

    How Plan Caching Reduces LLM Agent Costs

  10. 5 Jun 2026

    Stop filling your agent's context window just because you can

  11. 21 May 2026

    I benchmarked 5 embedding models across 4 datasets

  12. 19 May 2026

    Why reranking matters with cross-encoders

  13. 11 May 2026

    A beginner-friendly guide to the GGUF model format

  14. 27 Mar 2026

    Semantic Caching in Production

Learning Out Loud03
Selected noteReading preview

5 Oct 2026

Using Jev to verify cached LLM answers

I compared three ways of using Jev to check cached answers, then tested it against GPT-4.1 mini. The checks were cheaper and faster, with a trade-off in how many answers could be reused.

Read in the notebook →

13 Sept 2026

How Does the Dot Product Measure Similarity?

A step-by-step journey from vectors and the law of cosines to why the dot product of normalised vectors measures directional similarity.

Read in the notebook →

28 Aug 2026

How LLMs Turn Text into Embeddings

A first-principles look at how GPT-style language models turn raw text into token IDs, token embeddings, and positional embeddings.

Read in the notebook →

13 Aug 2026

Designing AI Agents from the Outside In

The PEAS framework offers a way to design AI agents from the outside in by defining success, environment, actions, and observations before writing code.

Read in the notebook →

30 Jul 2026

Could agents learn to work better from their own runs?

Most research on agent memory focuses on remembering the user. I built a browser agent that learned from its own runs, and the lessons that transferred surprised me.

Read in the notebook →

10 Jul 2026

Context engineering needs a context engine

Stronger models help, but agent reliability may depend just as much on better context systems around them.

Read in the notebook →

6 Jul 2026

Borrowed confidence is fragile in agentic systems

Production-like evals revealed the retrieval architecture I actually needed and reminded me that confidence in agentic systems has to be earned, not borrowed.

Read in the notebook →

19 Jun 2026

The lethal trifecta in AI agents

When agents can read private data, process untrusted content, and communicate outward, prompt injection becomes a much more serious security problem.

Read in the notebook →

17 Jun 2026

How Plan Caching Reduces LLM Agent Costs

Plan caching reuses planning templates across similar agent tasks, cutting cost and latency without throwing away accuracy.

Read in the notebook →

5 Jun 2026

Stop filling your agent's context window just because you can

Bigger context windows do not remove failure modes. They create new ones when we stop being intentional about what goes into an agent's context.

Read in the notebook →

21 May 2026

I benchmarked 5 embedding models across 4 datasets

I benchmarked five embedding models across four NanoBEIR datasets and found that bigger embeddings did not always produce better retrieval.

Read in the notebook →

19 May 2026

Why reranking matters with cross-encoders

Bi-encoders make retrieval fast, but cross-encoders expose why reranking matters when meaning depends on the query.

Read in the notebook →

11 May 2026

A beginner-friendly guide to the GGUF model format

GGUF made local LLM inference feel practical by packaging model weights, vocabulary, hyperparameters, and architecture metadata into one runnable format.

Read in the notebook →

27 Mar 2026

Semantic Caching in Production

Repeated user intents can quietly inflate LLM cost and latency. Semantic caching helps, but production use comes with trade-offs.

Read in the notebook →
Learning Out Loud04