A larger context window lets an application send more material to a model. It does not decide which material is authoritative, relevant, current, or accessible to this user. Those remain information-management decisions.

The choice is not simply “retrieval for small models, full context for large models.” A short curated source set may work well directly in context; a large permissioned corpus may still need filtering and retrieval. The right comparison is the task, evidence, and operating budget.

Capacity and use are different

Lost in the Middle showed that the models and tasks it studied could use long inputs unevenly depending on where relevant information appeared. This is a foundational finding, not proof that every later model has identical behavior. It motivates a local test of position, distractors, and evidence density.

LaRA, a 2025 benchmark comparing RAG and long-context approaches, also finds no simple universal winner. Its results depend on the tasks and systems evaluated. Use that evidence to design a comparison rather than to choose an architecture by slogan.

Define the evidence unit

For a synthetic policy question, the answer may require a general rule, a regional exception, and the effective date. Retrieving only the rule is incomplete even if that passage is relevant. Sending the full handbook does not guarantee that the model applies the exception.

Annotate the evidence needed for each evaluation task. Measure whether it enters the context and whether the answer preserves its conditions. Distinguish relevant-document recall from enough evidence to answer. Test queries that require no answer or a clarification.

Compare three realistic paths

One candidate can send a small approved source set directly. A second can use lexical or dense retrieval with permission filters and reranking. A third can use a hybrid: retrieve a document group, then provide richer surrounding context. Give each the same source versions and supported task.

Keep actual cost and latency visible. Full-context input may repeat large material across requests; retrieval adds indexing and selection work. Context caching can change the comparison, but its scope, expiry, source version, and permission behavior must be tested.

Compaction is a lossy transformation

Long-running agents accumulate tool results and conversation history. Anthropic's context-engineering guidance discusses curating that state, including compaction and selective retrieval. Treat a summary as a derived artifact that can omit a condition or distort a source.

Keep stable references to important evidence and unresolved decisions. Test whether compaction preserves the user's constraint, approval state, source version, and uncertainty. Do not summarize “approved to draft” into “approved to send.” An apparently minor loss can change the action boundary.

Test position and irrelevant information

Move necessary evidence to different positions, add a plausible conflicting draft, and vary the amount of unrelated content. Keep the legitimate task unchanged. These tests show how the chosen system handles its own operating context; they are not a complete robustness proof.

Separate a wrong answer caused by missing evidence from one caused by misusing present evidence. A retrieval fix addresses the first failure; generation or task framing may be needed for the second.

Make context a maintained contract

Record which sources, tool schemas, policies, and history enter a turn, along with their versions and access scope. Re-evaluate when the source universe or workflow changes. A context-window specification is a capacity constraint, not a claim that every token is equally useful.

The engineering objective is enough trustworthy information to complete the task at acceptable cost. Sometimes that means more context. Often it means better selection and clearer boundaries.

Sources and further reading