Optimizing RAG Accuracy with BGE-Reranker and FlagEmbedding

Artificial Intelligence tutorial - IT technology blog
Artificial Intelligence tutorial - IT technology blog

When Vector Search Returns Results That "Look Relevant but Miss the Point"

Most basic RAG pipelines follow a familiar workflow: receive a user query → generate query embeddings → run Cosine Similarity search on a Vector DB → pass the Top 5 chunks into the prompt for the LLM. This process works reliably during the demo phase with a few dozen PDF files.

Everything changes when moving to production. With an internal knowledge base spanning tens of thousands of pages, the Retrieval Hit Rate often drops drastically.

Consider a real-world scenario:

  • User query: "Are probationary employees covered by company health insurance?"
  • Top 3 Vector Search results: Returns 3 clauses covering insurance coverage for full-time permanent employees because of the high keyword density for "employees", "health insurance", and "coverage". Meanwhile, the specific exclusion clause for probationary employees drops all the way to the 9th position.

The outcome? The LLM receives the wrong context and confidently hallucinates that probationary employees are fully covered.

Simply increasing top_k to 15 or 20 does not solve the issue. You will encounter the Lost in the Middle phenomenon, where the LLM misses critical details buried in the middle of a large context window. Token costs and response latency also increase significantly.

Why Do Bi-Encoders Frequently Retrieve Irrelevant Context?

To resolve this, we must examine the algorithmic and hardware trade-offs of Dense Retrieval models (such as text-embedding-3-small or bge-base-en-v1.5).

These models utilize a Bi-Encoder architecture:

  1. Queries and Documents are processed independently through separate branches and compressed into fixed-size vectors (e.g., 768 or 1536 dimensions).
  2. The system measures relevance via Dot Product or Cosine Distance between the two vectors.

The primary limitation is that Query and Document tokens do not directly interact (no cross-attention). Compressing a 400–500 word passage into a single dense vector inevitably discards subtle conditional logic and negations.

3 Approaches to Improving Retrieval Accuracy

AI engineers typically evaluate three main strategies:

  • 1. Hybrid Search (BM25 + Dense Retrieval): Combines exact keyword matching with semantic vector search. This handles queries with specific error codes, product SKUs, or entity names well, but still struggles with complex conditional reasoning.
  • 2. LLM-based Filtering: Uses smaller models (such as Llama-3.1-8B or GPT-4o-mini) to scan and filter the Top 20 chunks. While highly accurate, this introduces significant latency (often 1–2 extra seconds) and substantially increases API costs.
  • 3. Two-stage Retrieval with Cross-Encoder Reranker (Recommended): Splits the search pipeline into two distinct phases:
    • Stage 1 (Fast Retrieval): Uses a Bi-Encoder to quickly scan millions of records in Qdrant or Milvus, gathering roughly 25–40 initial candidates (~10–20ms latency).
    • Stage 2 (Re-ranking): Uses a Cross-Encoder to compute full token-level cross-attention matrices between the query and each candidate, selecting the most accurate Top 3–5 chunks for the LLM.

Hands-on Implementation with BGE-Reranker and FlagEmbedding

BGE-Reranker by BAAI is among the most effective open-source Cross-Encoder model families available. Combined with the FlagEmbedding library, setup takes under 10 minutes.

1. Environment Setup

pip install FlagEmbedding torch

2. Quick Context-Sensitivity Test with FlagReranker

The script below demonstrates the model’s fine-grained contextual differentiation:

from FlagEmbedding import FlagReranker

# bge-reranker-v2-m3 provides excellent multilingual support, including Vietnamese
# use_fp16=True reduces VRAM usage by half and accelerates inference
reranker = FlagReranker('BAAI/bge-reranker-v2-m3', use_fp16=True)

query = "Are probationary employees entitled to company-sponsored health insurance?"
passages = [
    "All permanent employees signing contracts of 1 year or longer receive premium health insurance and mandatory statutory health coverage.",
    "During the 2-month probation period, employees are not eligible for company statutory health insurance; probationary base salary includes direct allowances.",
    "Medical reimbursement procedure: employees must submit official invoices to Accounting within 7 business days following discharge."
]

pairs = [[query, p] for p in passages]
scores = reranker.compute_score(pairs)

results = sorted(zip(passages, scores), key=lambda x: x[1], reverse=True)

print("=== RERANKING RESULTS ===")
for doc, score in results:
    print(f"Score: {score:.4f} | Content: {doc[:80]}...")

The results show that the second passage (containing the exact policy for probationary staff) receives a significantly higher score (~5.82 vs. -2.14 for the first passage), despite the first passage having a higher raw keyword overlap.

3. Integrating the Reranker into a Retriever Class

Here is a boilerplate class for your RAG backend:

from typing import List, Dict, Any
from FlagEmbedding import FlagReranker

class TwoStageRetriever:
    def __init__(self, vector_store: Any, model_name: str = 'BAAI/bge-reranker-v2-m3'):
        self.vector_store = vector_store
        self.reranker = FlagReranker(model_name, use_fp16=True)

    def retrieve_and_rerank(
        self, 
        query: str, 
        initial_top_k: int = 30, 
        final_top_k: int = 4
    ) -> List[Dict[str, Any]]:
        # Step 1: Rapid candidate retrieval via Vector DB
        initial_docs = self.vector_store.similarity_search(query, k=initial_top_k)
        if not initial_docs:
            return []

        # Step 2: Build query-document pairs
        doc_texts = [doc.page_content for doc in initial_docs]
        pairs = [[query, text] for text in doc_texts]
        
        # Step 3: Compute cross-attention relevance scores
        scores = self.reranker.compute_score(pairs)
        
        # Step 4: Attach scores and sort to get Top-K
        scored_docs = []
        for doc, score in zip(initial_docs, scores):
            doc.metadata["rerank_score"] = float(score)
            scored_docs.append(doc)
            
        scored_docs.sort(key=lambda x: x.metadata["rerank_score"], reverse=True)
        return scored_docs[:final_top_k]

4. Production Deployment Best Practices

  • Selecting the Right Model Size:
    • bge-reranker-base (~280M params, ~600MB VRAM footprint): Best suited for ultra-low latency targets (<25ms) or CPU deployment clusters.
    • bge-reranker-v2-m3 (560M params, ~1.2GB VRAM with FP16): Optimal choice for Vietnamese and broader multilingual retrieval contexts.
  • Measuring Latency: On standard T4 GPUs, reranking 30 chunks with bge-reranker-v2-m3 (FP16) takes roughly 45–60ms. This trade-off is well worth the gain in Hit Rate@3, moving from ~68% to over 91%.
  • Optimal Candidate Batch Size: Keep initial_top_k between 25 and 40. Pushing beyond 80 chunks increases latency linearly with diminishing accuracy returns.
Share: