Can Long Context Replace RAG? Choose by Evidence Scope and Cost per Qualified Answer

Longer model inputs give enterprise knowledge systems another option: reading a defined document collection directly. Capacity alone does not ensure complete evidence or lower cost. This article compares retrieval, long-context and hybrid approaches through evidence scope, workload and end-to-end cost, then sets out a practical comparative evaluation.

Choosing an enterprise knowledge architecture should start with a concrete question: which materials must an answer consider, and how can the team detect a missing exception that changes the conclusion? Long context changes how materials reach the model; it does not define their scope. Compare approaches on the same real tasks, permissions, document versions and acceptance criteria before deciding what to deploy.

The change concerns reading scope, not the disappearance of knowledge engineering

Retrieval-augmented generation (RAG) selects a limited set of relevant material before a model composes an answer. A long-context approach supplies a larger, predefined document collection directly. These are compatible choices: a system can retrieve whole documents and then read their contents and attachments with a long-context model. The architectural question is which evidence-selection decisions belong to retrieval and which belong to the model.

Li and colleagues compared the approaches on specified models and public datasets in 2024. With sufficient resources, long context performed better on average, while RAG retained a cost advantage; the paper also proposed hybrid routing.[1] These findings do not establish a purchasing verdict for every current model or validate an entire enterprise corpus. Here they motivate comparative baselines, rather than transferable performance numbers.

There is counterevidence to equating capacity with dependable use. The 2024 Lost in the Middle study demonstrated effects from evidence position in particular tasks. Google's long-context documentation also distinguishes finding one item from extracting multiple items.[2][3] Input capacity, single-item retrieval performance and complete business-answer quality therefore need separate evaluation.

Define the smallest complete evidence scope first

Our proposed engineering method begins with an evidence checklist for each question type: required sources, amendments with higher precedence, whether every relevant entity must be covered, and whether absence from the results can establish actual absence. This is a proposed method, not an industry standard. It shifts the discussion from input length to the consequences of omissions.

A maintenance instruction may depend on a few pages of a particular manual version. Comparing a purchase agreement with supplementary terms requires reading potentially overriding provisions together. Listing every obligation for a project requires a defined project-document universe and a completeness check. All three look like document questions, but they have different omission risks.

The earlier discussion of permission-aware enterprise RAG explains access boundaries and evidence applicability. Those requirements also apply to long context. Spare input capacity does not authorize confidential attachments, and outdated versions should not be mixed in for the model to resolve unaided. Apply authorization and version selection when building the input.

Test different starting architectures for different workloads

For local factual questions over a large corpus, start with retrieval. It limits input to candidate evidence and suits questions whose support can be located through keywords or semantic similarity. Its costs include chunking, index maintenance and recall evaluation. When exceptions are dispersed, retrieving the main clause can still produce a wrong answer. Increasing candidate counts may reduce omissions without establishing completeness.

For an explicitly selected file collection, such as a review package with fixed attachments, direct long-context input is a useful baseline. It removes some fragment-selection work and supports cross-section comparison. The package still needs parsing, version and access checks, with input room reserved for instructions, the question and the response. Do not silently truncate an oversized package; narrow the task explicitly or use staged reading.

When the corpus is larger but relevant documents can be identified first, test document retrieval followed by full-document reading. Retrieval chooses files; long context handles their internal relationships. Document-selection omissions remain possible, but important attachments may no longer be fragmented. When a question requires summing all transactions or exhaustively enumerating records, prefer controlled database queries and deterministic computation. Text retrieval cannot itself prove complete enumeration.

Compare end-to-end cost per qualified answer

We propose dividing total operating cost over the same evaluation period by the number of qualified answers. Include document parsing and updates, indexing and retrieval, model input and output, failed attempts and necessary human review. Show development and migration costs separately so one-time spending does not obscure ongoing differences. An answer qualifies only if it meets factual, evidential, authorization and response-time requirements—not merely because text was returned.

Long context may remove some indexing work while increasing repeated input and waiting time. RAG can reduce input volume but requires continuing retrieval-quality maintenance. For repeated questions over the same materials, evaluate context caching. Google's documentation describes caching as a way to reduce repeated-context cost and notes that longer inputs generally increase time to first output.[3] Actual savings depend on reuse, expiry and the chosen service's billing terms.

Measure cold-cache, warm-cache and post-update requests separately; do not assume every request hits a cache. Alongside median latency, report slower-request percentiles and timeout frequency. Report total requests and qualified answers together with unit cost. A cheap route that often requires correction may simply transfer expense to the business team.

Use a constructed case to expose omissions

The following is a constructed scenario, not an IDENIFE project or measured result. A team reviews an equipment purchase package containing a main contract, an acceptance attachment and an amendment. It asks whether on-site training is mandatory for this delivery. The contract gives a general requirement, the attachment defines equipment scope, and the amendment changes arrangements for this delivery. Reading only the main contract could yield a fluent but incorrect answer.

Create three baselines over one authorized document snapshot: fragment retrieval, the complete file collection, and document retrieval followed by full reading. Reviewers first identify the necessary clauses and precedence relationships. Then assess whether each approach finds the evidence, resolves overrides correctly and identifies missing materials. Do not treat an answer generated by one candidate system as the reference answer.

Add a similar but out-of-scope old contract, move the decisive amendment within the input, and remove one necessary attachment. Position changes test stability; the old contract tests version boundaries. When an attachment is missing, a qualified response explains why the conclusion cannot be determined instead of repeating a previous answer. Include simple questions that do not need full reading, so a few demanding cases do not dictate the cost of every request.

Hybrid routing needs explainable escalation rules

Low confidence does not automatically justify resending the full corpus. First record whether document scope is known, required sources are present and input fits the budget. Use those facts to decide whether to expand reading. Model-reported confidence can help, but should not be the sole release condition. Research on hybrid routing is relevant; enterprise routing also needs access, latency and evidence constraints.[1]

Keep the original question, material versions, attempted route and newly added evidence during escalation. Otherwise a retry can become an unexplained alternative answer. Broader input must not bypass authorization, and genuinely missing sources call for additional material or a limited conclusion. Operations teams can change routes by question type without permanently committing the entire knowledge system to one architecture.

Acceptance should produce three outputs: quality and omission records by task class, a cost account including maintenance and review, and explicit behavior at failure boundaries. Increase traffic only where an approach meets those conditions for its workload. Whether long context replaces RAG should remain a repeatable engineering decision, rather than a brand choice inferred from window size.

References

  1. [1] Li et al. (EMNLP 2024): Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
  2. [2] Liu et al. (TACL 2024): Lost in the Middle: How Language Models Use Long Contexts
  3. [3] Google: Long context — Gemini API
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.