The lost in the middle effect is the tendency for language models to use information at the start or end of a long context more reliably than information in the middle. It is an evaluation finding that exposes a positioning bias, not a failure of retrieval alone.
How the lost in the middle effect appears in long-context evaluation
The lost in the middle effect shows up when a model can answer correctly from the opening or closing parts of a long prompt, yet becomes less reliable when the same evidence is placed in the middle. That makes it a positioning bias, not just a retrieval problem, because the information is present but not equally used.
In practice, this effect matters most in long-context tasks such as document QA, policy review, incident analysis, and tool-augmented prompts where the relevant fact may sit far from the most salient context windows. It is a reminder that context length alone does not guarantee effective context use.
Researchers often treat the effect as a probe of attention allocation and context utilisation, not as proof that the model cannot retrieve the answer at all. A model may still be capable of finding the information, but fail more often when the supporting text is buried in the middle of a long sequence.
Why it matters for prompting and evaluation
The lost in the middle effect changes how practitioners should interpret long-context performance. A strong score on short or front-loaded prompts may overstate robustness if the model has not been tested against mid-context evidence, distractors, and reordered inputs.
For evaluation, it is useful to vary where the critical evidence appears, whether the answer is repeated, and how much surrounding noise is present. The goal is to measure whether the system can use context across the full prompt, not just at the boundaries where models tend to do better.
The effect is also relevant to prompt design. If a task depends on one or two critical facts, placing them only once in the middle of a very long context can make performance look worse than it should, even when the model is otherwise competent. That can mislead teams into blaming the retrieval layer when the issue is broader context utilisation.
For a broader view of long-context governance and downstream identity-security consequences in operational systems, NHI programmes often see the same pattern of hidden dependence when key details are not surfaced early enough in workflows, as reflected in NHI Mgmt Group’s Ultimate Guide to NHIs.
Common failure modes and practical interpretation
The core failure mode is uneven attention across the prompt, but several related issues can make it look worse. Long prompts may contain distractors, repeated instructions, or multiple candidate facts, and the model may overweight the earliest or latest cues even when the middle contains the correct evidence.
This means the effect is best read as a sign of sensitivity to layout, salience, and ordering. It does not automatically mean the model is weak at reasoning, nor does it prove that the retrieval system failed to supply the right material. A system can retrieve correctly and still underuse the retrieved evidence.
That distinction matters operationally. If the middle of the context is where authoritative data, policy exceptions, or control details are being placed, performance degradation can be caused by prompt architecture rather than data quality.
What practitioners should test and watch for
Teams evaluating long-context models should check whether answer quality changes when the same evidence is moved to the beginning, middle, and end of the prompt. They should also test with competing distractors and with repeated facts, because those patterns reveal whether the model is reading for position or for substance.
What to watch for: unstable answers across reordered prompts, improved performance when critical evidence is duplicated near the end, and unexpected failures on documents that are otherwise well within the model’s context window. Those signals usually indicate a context-use issue, not just a retrieval miss.
Practitioner takeaway: Treat long-context evaluation as an ordering test as much as a knowledge test; if the answer only appears when the evidence is near the edges, the model is not yet using the full prompt reliably.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 address the attack and risk surface, while NIST CSF 2.0 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.2 — Cybersecurity Risk Management Strategy | Long-context positioning bias affects governance of AI-enabled security workflows. |
| Recommendation — Assess long-context model limitations in governance reviews before relying on them for operational decisions. | ||
| NIST AI RMF | GOVERN — Govern AI Risk | The effect is an AI reliability issue that should be governed as a model-risk concern. |
| Recommendation — Document context-length failure modes in AI risk governance and validation. | ||
| OWASP Agentic AI Top 10 | LLM-01 — Prompt Injection and Instruction Hierarchy | Prompt ordering and instruction placement directly affect how an AI system follows context. |
| Recommendation — Place critical instructions redundantly and test for instruction adherence across long prompts. | ||
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 23, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org