~/writing
Modern AI Systems

Hybrid Search, Re-ranking, RAG and RAG Evaluation

Notes on hybrid search, reranking with RRF, retrieval methods like BM25 and multi-query retrieval, and RAGAS metrics for evaluating RAG systems.

As seen previously, there are trade-offs between choosing just vector search or just keyword search. In order to get the best of both worlds, modern systems use a hybrid search, which combines various types of searches to produce multiple lists of relevant documents. These combined results are then passed to a reranking module, which optimizes the final ranking order based on relevance signals or learned scoring.

Other retrieval methods include Best Matching 25 (BM25), multi-query retrieval and ensemble.

BM25: Used for keyword search, ranks based on term frequency, document length normalization and inverse document frequency.

BM25 counts how often a query word appears in a document but using a non-linear curve to prevent long documents filled with repeated keywords from dominating results, through diminishing returns.

It also adjusts scores based on document length, so a short document containing a query word is ranked higher than a long document containing the same query word since the shorter document is likely more focused and relevant.

Lastly, BM25 weights rare query words much higher than common ones. A rarer term like “macronutrients” would rank higher than a common term like “food”.

Multi-query retrieval: Retriever uses a LLM to produce multiple versions/interpretations of the same query, and the search outcomes are aggregated to give a more comprehensive final query response. Useful for product recommendation systems, document retrievals and academic research.

Ensemble: Ensemble retriever combines keyword search (like BM25) and dense retrievers (like vector search) to produce a list of relevant documents. Uses methods like reciprocal rank fusion (RRF) to combine scores from multiple retrieval methods to produce a final ranking. One example would be the hybrid search mentioned above.

Hybrid search and reranking pipeline

Since hybrid search was used, there would be two separate ranking lists: one for keyword search, and the other for vector search. The scores from each list are then combined through RRF and reranked before being returned to the LLM.

RRF

When a user fires off a query, both keyword searches and vector searches are made. The reciprocal rank score is calculated using the formula below:

Score = 1/ (rank + K)

where K decides the sensitivity to rank positions and rank is the position of the document in the list.

In some implementations, the hyperparameter alpha can be fine-tuned between 0-1, where a value closer to 0 favors keyword search, and a value closer to 1 favors vector search. This can be tweaked depending on your use case.

RRF scoring example combining sparse and dense search

Sparse search = keyword search, Dense search = vector search

RAGAS metrics

5 main metrics:

  • Answer relevancy: How relevant the generated response is to the given input
  • Faithfulness: Whether the generated response contains hallucinations to the retrieval context
  • Contextual Relevancy: How relevant the retrieval context is to the input
  • Contextual Recall: Whether the retrieval context contains all the information required to produce the ideal output (for a given input)
  • Contextual Precision: Whether the retrieval context is ranked in the correct order (higher relevancy goes first) for a given input

The quality of RAG responses are usually determined by the retrieval and generator. When retriever fails, it’s usually due to:

  • Poor chunking strategy
  • Uninformative embeddings
  • Weak reranking logic
  • Suboptimal top-k setting

When generator fails, it’s usually due to:

  • Ignoring key information
  • Focusing on the wrong details
  • Misreading prompt structure
  • Weak prompts/model limitations

Contextual Relevancy

This quantifies the proportion of retrieved text chunks that are relevant to the input and is a measure of how well your top-k and chunk size is configured. Smaller chunks can give you more precise information but might require retrieving more of them (higher top-K) to cover enough context, while larger chunks may capture broader meaning but risk including irrelevant details.

Contextual Recall

The whole purpose of contextual recall is to assess whether the retrieval context contains all the necessary information to produce the ideal output. Without contextual recall, anyone can achieve perfect contextual relevancy by retrieving as little information as possible.

Contextual Precision

The purpose of precision is to make sure that the context you are feeding into your generator is not just relevant, contains all the information, but also in the correct order for an LLM to consider each text chunk’s importance appropriately.

Answer Relevancy

The answer relevancy metric quantifies the proportion of the generated output that is relevant to the given input. It is extremely straightforward and a direct measure of how well your model can follow instructions in your prompt template.

Faithfulness

The faithfulness metric quantifies the proportion of undisputed, truthful facts in the retrieval context that were not contradicted in the generated output. Essentially, it measures the hallucination rate of your LLM.

Custom Metrics

Standard RAG metrics evaluate basic mechanical accuracy (retrieval and generation) but fail to cover application-specific business requirements. Custom metrics tailor evaluation to real-world deployment needs:

Domain compliance and guardrails: Enforces regulatory rules and safety constraints (preventing unauthorized financial, legal, or medical advice)

Brand persona, tone: Measures alignment with required communication styles (maintaining empathy, professional voice, or avoiding technical jargon)

Formatting/Schema: Verifies that model outputs strictly adhere to required structures (JSON schemas, specific bulleted layouts, or precise section counts)

Fallback behaviors: Evaluates how well the model handles missing information (outputting an explicit “I don’t know” rather than guessing or outputting generic filler)