~/writing
Modern AI Systems

RAG vs CAG vs MAG

Comparing RAG, CAG, and MAG — how each approach retrieves and grounds context for LLMs, and where each one fits.

Retrieval-augmented generation (RAG), Cache-augmented generation (CAG) and Memory-augmented generation (MAG) all solve the same underlying problem: knowledge that LLMs refer to is static and bounded by its context window, so it needs a way to bring in information that it wasn’t trained on or cannot remember. The main difference in these approaches is the where the extra information comes from and how its delivered.

RAG

Mentioned in previous articles, basically pulls relevant information from an external knowledge source like vector databases or document stores at the time of the query

Strengths of RAG:

  • Real time freshness

Data updates can be done in real time, and the updates will propagate to answers almost immediately without retraining

  • Fewer hallucinations

Grounding LLM responses in retrieved documents gives the model a factual anchor and reduces hallucination

  • Flexible sourcing

Knowledge can be drawn from structured databases, unstructured text, pipelines can be tailored for different use cases etc

Limitations of RAG:

  • Latency

Retrieving information from an external database every query adds computational time

  • Retrieval quality dependency

The quality of responses is highly dependent on the quality of retrieval, as irrelevant chunks degrade the response

CAG

CAG became increasingly practical as context windows grew into the hundreds of thousands to millions of tokens. It relies on two mechanisms:

  1. Knowledge caching

This is where reference material is preloaded directly into the model’s extended context to reuse across multiple queries without external fetches from knowledge sources

  1. KV (key-value) caching

Attention states are computed while processing tokens are stored so that they can be reused for similar/repeated queries without being recomputed

The difference between RAG and CAG can also be seen as the difference between looking up answers in a reference book as opposed to already having a cheat sheet that has been prepared to refer to.

Strengths of CAG:

  • Speed

Reusing cached computation cuts response times especially for repeated queries

  • Consistency

Stored context avoids response drift throughout a session, which is useful for chatbots and workflow automation

Limitations of CAG:

  • Staleness

Cached knowledge does not reflect changes made after the cache was built, requires frequent cache updates

  • Memory cost

Large caches require significant compute/RAM

MAG

Instead of injecting external knowledge, MAG allows the model to recall past interactions, prompts, responses and even users’ stated preferences. This creates continuity and personalization even though the underlying model is stateless between calls. MAG is generally good when personalization is needed, such as a recommendation engine remembering past preferences of users, or for continuity during long-running interactions where the model needs to behave consistently rather than restart context each time.

General use cases

  • Use RAG when there is frequently changing information, large/diverse knowledge bases or if stale answers are costlier compared to some latency
  • Use CAG when knowledge is stable (won’t change often), or when there is a high volume of repetitive queries and where latency is important
  • Use MAG when you need the system to feel consistent and personalized across a session or over time regardless of whether the knowledge is fresh or stable.