Skip to content

Context Window Growth Explained

Cheatsheet timeline of context window growth: from GPT-2's 1K tokens to the million-token era — the milestones that changed what prompts could hold.

In six years the usable context grew a thousandfold: from 1K tokens (a page) to 1M+ (a codebase). Each jump changed the architecture around the model — from clever summarization to just pasting everything.

Reference table · 10 entries
10 of 10 rows
The 2K era — prompt engineering is born
2019GPT-21KA paragraph of conditioning; few-shot barely fits.
2020GPT-32KFew-shot prompting works — the pattern-in/pattern-out era.
The 8K–32K era — documents fit
2022GPT-3.5 / early GPT-44KChat with history; whole emails.
2023GPT-4 32K · Claude 1 100K32K–100KWhole documents and small codebases — 'just paste it' begins; RAG's first challenger.
The 128K–1M era — the corpus fits
2023–24GPT-4 Turbo128KBook-length context as a standard tier; needle-in-haystack benchmarks go mainstream.
2024Gemini 1.5 Pro1MHour-long video, whole repos — retrieval-optional for many tasks.
2025+Frontier models200K–1M+Long context as table stakes; cache pricing makes giant prompts affordable.
The caveats that survived every jump
—lost in the middle—Recall is U-shaped: start and end of context beat the middle.
—cost & latency—Attention cost grows with length; long context bills per token.
—prompt caching—Stable prefixes cache at ~10% — architecture matters again (see the Cache Breakpoint Planner).