Building a Retrieval-Augmented Generation (RAG) system is not just about connecting a large language model to a vector database. The true bottleneck for accuracy lies in the retrieval phase. If the retriever brings irrelevant context to the LLM, the final answer will be flawed regardless of the model's intelligence. To achieve high-precision RAG, we must move beyond simple keyword or semantic matching by combining the strengths of multiple retrieval strategies and refining the results with re-ranking.
The Limitations of Single-Strategy Retrieval
Most initial RAG implementations rely on either sparse retrieval (like BM25) or dense retrieval (using embeddings like Sentence-BERT). Each has distinct blind spots.
- Sparse Retrieval (BM25): Excels at exact keyword matching and understanding specific entity names, codes, or technical terms. However, it fails to capture semantic similarity or paraphrasing.
- Dense Retrieval (Embeddings): Captures semantic meaning and context well, allowing it to find documents that discuss the same concept but use different words. However, it often struggles with unique identifiers, specific product SKUs, or rare technical jargon.
By using only one, you sacrifice either precision on specific terms or recall on semantic variations. Hybrid search combines both signals to mitigate these weaknesses.
Building a Hybrid Retriever in Haystack
The Haystack framework provides an elegant abstraction for combining retrievers. You can orchestrate a BM25Retriever and an EmbeddingRetriever in parallel and then combine their results.
from haystack import Pipeline
from haystack.components.retrievers.in_memory import BM25Retriever, EmbeddingRetriever
from haystack.components.routers import SwitchRouter
from haystack.components.joiners import DocumentJoiner
from haystack.components.builders import DocumentJoiner
# Initialize Retriever Components
# Note: In a real pipeline, you would connect these to your InMemoryDocumentStore
bm25_retriever = BM25Retriever(document_store, top_k=20)
embedding_retriever = EmbeddingRetriever(document_store, top_k=20)
# Initialize the Joiner
# This component merges documents from both retrievers
# The score_type determines how scores are combined (e.g., 'distributional', 'arithmetic')
doc_joiner = DocumentJoiner(
join_mode="concatenate",
duplicate_documents="skip",
score_normalization="min_max",
score_threshold=0.3
)
# Construct the Hybrid Search Pipeline
pipeline = Pipeline()
pipeline.add_component("bm25", bm25_retriever)
pipeline.add_component("embedding", embedding_retriever)
pipeline.add_component("joiner", doc_joiner)
# Connect the components
# Both retrievers run in parallel, and their outputs are fed into the joiner
pipeline.connect("bm25", "joiner")
pipeline.connect("embedding", "joiner")
In this configuration, the pipeline queries both indexes simultaneously. The DocumentJoiner then merges the lists of documents. Crucially, you should normalize the scores because BM25 scores (which are unbounded and query-dependent) and Cosine Similarity scores (bounded between -1 and 1) are not directly comparable. Using score_normalization="min_max" scales both sets of scores to a common range, allowing for fair competition during the joining process.
The Critical Role of Re-Ranking
While hybrid search improves recall, the top 20 or 30 documents retrieved are still a "bag of results." Some may be only marginally relevant. Feeding all of these to the LLM wastes context window space and can introduce noise. This is where re-ranking enters the picture.
Re-rankers use Cross-Encoder models. Unlike Bi-Encoders used in vector retrieval (which encode query and document independently), Cross-Encoders process the query and the document together in a single attention head. This allows the model to perform deep, token-level interaction, providing a significantly more accurate relevance score.
Integrating a Re-Ranker into the Pipeline
We append a SentenceTransformersReRanker or a dedicated API-based re-ranker (like Cohere or Jina) to the end of the retrieval stage. The re-ranker takes the combined top-N documents from the hybrid search and re-orders them, selecting only the top-K most relevant ones for the generator.
from haystack.components.rerankers import TransformersSimilarityRanker
# Initialize Re-Ranker
# Use a strong cross-encoder model for best performance
re_ranker = TransformersSimilarityRanker(
model="cross-encoder/ms-marco-MiniLM-L-6-v2",
top_k=5,
batch_size=16
)
# Update the Pipeline
pipeline.add_component("re_ranker", re_ranker)
# Connect the joiner to the re-ranker
pipeline.connect("joiner", "re_ranker")
By setting top_k=5 in the re-ranker, we drastically reduce the context passed to the LLM. Instead of sending 40 potentially noisy documents, we send only the 5 documents that the cross-encoder has verified as highly relevant. This typically results in a measurable improvement in answer precision and a reduction in latency due to smaller prompt sizes.
Practical Considerations and Optimization
Implementing this architecture requires balancing several factors:
- Chunking Strategy: Ensure your documents are chunked appropriately before indexing. If chunks are too large, the re-ranker may struggle to find the specific relevant sentence. If too small, you lose context. A range of 200-500 tokens with 10-20% overlap is a good starting point.
- Scoring Thresholds: Use the output of the re-ranker to filter out weak matches. If the top score after re-ranking is below a certain threshold (e.g., 0.4 for normalized scores), it is often better for the system to admit "I don't know" than to hallucinate an answer based on weak evidence.
- Model Selection: For production environments, consider using hosted re-ranker APIs (such as Cohere's Rerank or Jina AI's Reranker) to offload the computational cost of cross-encoding from your local infrastructure, especially for high-throughput systems.
Conclusion
High-precision RAG is not a magic trick; it is the result of engineering a robust retrieval pipeline. By leveraging Haystack's component-based architecture, you can seamlessly integrate hybrid search to capture both semantic and lexical signals, followed by a cross-encoder re-ranker to ensure only the most pertinent context reaches the LLM. This multi-stage approach significantly reduces hallucinations and improves the factual accuracy of your AI applications, turning a basic chatbot into a reliable enterprise-grade knowledge assistant.