Beyond naive RAG: Make it contextual and Hybrid

Posted by Venkatesh Subramanian on October 04, 2026 · 4 mins read

In the past I’ve written about simple RAG strategies. However, this does not always work. Let’s break down how naive RAG works.

  1. Extract text from unstructured enterprise documents such as PDFs.
  2. Use either a fixed size or some logical paragraph based chunking.
  3. Embed the chunks into vectors, and store in a vector DB that enables ANN search.

Now imagine a MSA document that has :

┌─────────────────────────────────────────────────────────────────────────────┐
│ MASTER SERVICES AGREEMENT (50,000 tokens)                                   │
│ Page 1: "This Agreement is entered into by Acme Corp ('Customer') and       │
│          Beta Systems LLC ('Vendor') on January 15, 2026..."                │
│ ...                                                                         │
│ Page 42 (Chunk 84): "The Vendor shall maintain commercial general liability │
│                     insurance of not less than $5,000,000 per occurrence.   │
│                     Proof of coverage must be delivered within 30 days..."   │
└─────────────────────────────────────────────────────────────────────────────┘

When a user asks “What is Acme Corp’s required vendor liability insurance limit?”, the cosine similarity between the query and Chunk 84 is compromised because the entity relationship (Acme Corp -> Customer, Beta Systems -> Vendor) was established 41 pages earlier. Chunk 84 embedding will have no idea of what is the vendor name! This symptom is semantic starvation .

The Contextual Retrieval pattern

To solve semantic starvation without indexing entire multi-page documents as unmanageable mega-chunks, systems architects implement Contextual Retrieval.

Instead of embedding raw chunks in isolation, the ingestion pipeline utilizes a high-throughput, low-latency Small Language Model (SLM) to synthesize 50–100 tokens of contextual metadata situating the chunk within the parent document. This synthesized context is prepended directly to the chunk text prior to generating vector embeddings and building the inverted lexical index.

                      Parent Document (e.g., 50,000 Tokens)
                                     │
                                     ▼
                    Context Synthesizer (Fast SLM / LLM)
                  (KV / Prefix Cached on Parent Document)
                                     │
                                     ▼
                     Synthesized Explanatory Context:
    "This chunk is from the Master Services Agreement between Acme Corp 
     and Beta Systems LLC (Jan 2026), defining the Vendor's commercial 
     general liability insurance coverage thresholds."
                                     │
                                     ▼
                Prepend to Raw Chunk & Dual-Index
                                     │
          ┌──────────────────────────┴──────────────────────────┐
          ▼                                                     ▼
 [Dense Vector Index]                                  [Sparse BM25 Index]
  Vector space encodes global entity                    Inverted postings list includes:
  relationships & operational scope                     "acme", "corp", "beta", "2026"

Ingestion Cost Optimization: The Role of Prefix / KV Caching

While contextual enrichment eliminates retrieval blind spots, naive implementations create a computational and financial bottleneck during document ingestion.

The Naive Ingestion Multiplier

Consider an enterprise corpus of 5,000 regulatory filings or legal agreements, averaging 40,000 tokens per document. Segmented into 400-token units, each document yields approximately 100 chunks.

  • For every chunk, the context synthesizer must ingest the entire 40,000-token parent document to generate accurate contextual metadata.
  • Naive computation:
    100 text{ chunks} times 40,000 text{ input tokens} = 4,000,000 text{ input tokens per document}$$
  • For a corpus of 5,000 documents:
    5,000 times 4,000,000 = 20,000,000,000 text{ (20 Billion) input tokens}

At standard API or on-prem inference costs, naive contextual ingestion is cost-prohibitive.

The Systems Solution: Prefix / KV Caching

Modern inference engines (whether cloud managed APIs or self-hosted engines like vLLM, SGLang, and TensorRT-LLM using RadixAttention) support KV (Key-Value) Prefix Caching.

Because the parent document remains constant while only the individual chunk target changes per request, the attention key-value states for the parent document can be computed once and reused across all subsequent chunk synthesis calls.

Call 1 (Chunk 1):   [Parent Doc: 40k tokens (KV Cache Write)] ──► Full Compute / Base Cost
Call 2 (Chunk 2):   [Parent Doc: 40k tokens (KV Cache Read)]  ──► ~80-90% Latency & Cost Reduction
Call 3 (Chunk 3):   [Parent Doc: 40k tokens (KV Cache Read)]  ──► ~80-90% Latency & Cost Reduction
...
Call 100 (Chunk 100): [Parent Doc: 40k tokens (KV Cache Read)] ──► ~80-90% Latency & Cost Reduction

Summary

Hybrid retrieval using both sparse BM25 index and dense vector embeddings, small language model to prepend metadata for chunks, and caching for parent document metadata all result in a resource efficient and accurate RAG retrieval, that is also conntextually correct.


Subscribe

* indicates required

Intuit Mailchimp