Writing
Agent Evaluation
A breakdown of how AI agent evaluation differs from LLM evaluation, covering outcome, trajectory, safety, operational, and consistency metrics, plus how golden datasets and LLM-as-judge grading fit in.
Harness Engineering vs Loop Engineering
How agentic loops are structured, and why the harness — everything wrapping the model — is what makes an agent reliable in production.
Agent Memory
How AI agents store and retrieve memory — short-term vs long-term, semantic vs episodic vs procedural, and the storage systems behind them.
Workflows vs Agents
How workflows differ from agents, and the core building blocks — memory, tools, and planning — behind agentic systems.
RAG vs CAG vs MAG
Comparing RAG, CAG, and MAG — how each approach retrieves and grounds context for LLMs, and where each one fits.
Hybrid Search, Re-ranking, RAG and RAG Evaluation
Notes on hybrid search, reranking with RRF, retrieval methods like BM25 and multi-query retrieval, and RAGAS metrics for evaluating RAG systems.
Keyword Search vs Vector Search
Comparing how keyword search and vector search find results — exact word matching versus semantic similarity — and when to use each.
What Are Embeddings?
How embeddings turn objects into vectors machine learning models can compare, search, and reason over.