Skip to content Skip to footer

How Precomputed Context Can Make RAG Agents Faster and Cheaper

What Happened

Elastic described a context engine that prepares reusable knowledge about schemas, fields, entities and documents before an agent receives a question. It stores these “Knowledge Indicators” in an AI index and retrieves them through lexical, semantic, hybrid and agentic search. Agent traces feed a loop for improving the indicators over time [1].

In a simulated account-analysis task using four Elasticsearch indices and roughly 50 contracts, Elastic generated about 80 indicators. With the same task, prompt, data and model, its context-backed agent made 9 tool calls rather than 33, used 653,511 input tokens rather than 2,279,253, and ran in 64 seconds rather than 213 seconds. It also found a parent-company relationship the baseline missed. These are results from one evaluation, not a general performance guarantee [1].

Why It Matters to Businesses

Conventional retrieval-augmented generation (RAG) often makes agents rediscover database structure and business relationships on every request. Precomputing that context can move repetitive work out of the user-facing path, reducing latency and model spend. It does not eliminate live retrieval: an agent still needs current records and evidence to verify an answer [1].

For teams using LlamaIndex, LangChain or Haystack with Weaviate, Pinecone, Qdrant, Milvus, Elasticsearch or Vespa, the transferable lesson is architectural: treat organizational context as a maintained retrieval asset, not merely as more text in a prompt. The reported measurements do not establish a performance comparison among those products [1].

Kimbodo Engineering Perspective

Precompute stable knowledge; fetch volatile facts live. Schema descriptions and entity relationships may be good candidates for reusable context, while balances, permissions and contract status need checks against authoritative systems. A compact indicator is valuable only if the agent can trace it to its source and determine whether it is still valid.

We would compare this design with a strong baseline using the same questions, model and source data. The decision metric is total operating cost and answer quality—not token reduction alone—because extraction, indexing, refresh and evaluation also consume resources.

How We Would Implement It

  • Ingest approved database metadata and documents; assign each item a source identifier, owner, access scope and update timestamp.
  • Extract concise schema and entity indicators offline, link them to source records, and review a sample for accuracy before indexing.
  • Use lexical and vector retrieval together where exact names and semantic matches both matter. Return indicators as orientation, then query live systems for answer evidence.
  • Connect the retrieval layer to the agent framework through a narrow, authenticated tool interface. Log retrieval choices, verification steps, latency, tokens and corrections.
  • Refresh indicators when sources change; use failed searches and agent traces to prioritize improvements. Test against direct-retrieval baselines before rollout.

Risks, Costs and Security

Stale indicators can make an agent confidently follow the wrong relationship; broad retrieval can expose information across tenants or roles. Enforce source permissions at retrieval time, filter results before they enter the model context, and retain citations for verification. Treat retrieved text as untrusted data, not instructions.

Budget for extraction jobs, index storage, embedding or search infrastructure, monitoring and refresh. Measure whether those costs are outweighed by fewer tool calls and lower query-time token use in the business’s own workload [1].

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our RAG Development Services practice, or Estimate My RAG System.

Sources

  1. [1] How we’re building a context engine to cut agent token use by 71%

Leave a comment

0.0/5