All posts
// / Blog

Context window length is the most misunderstood feature in LLMs.

"Our model supports 200K tokens!" Great. But can it actually USE 200K tokens effectively?

I ran an experiment: buried a specific fact at different positions in a long context and asked the model to retrieve it. The results were revealing.

At the beginning: 98% accuracy. At the end: 95% accuracy. In the middle: 71% accuracy.

This is the "lost in the middle" problem. Most long-context models are significantly worse at using information from the middle of their context window.

What this means for RAG: don't just stuff your retrieved documents into the context. ORDER matters. Put the most relevant documents at the beginning. Important instructions at the beginning or end. Never rely on information buried in the middle of a long context.

For agent applications: don't assume the agent can track a long history of actions. Summarize periodically. Keep the current plan and recent actions at the beginning of the context.

The models are getting better at this. But "supports 200K tokens" doesn't mean "equally attentive to all 200K tokens."

Test your specific use case at your specific context lengths. Don't trust the marketing.

#LLM#ContextWindow#RAG#AIEngineering#MachineLearning#NLP