| The 2K era — prompt engineering is born | |||
|---|---|---|---|
| 2019 | GPT-2 | 1K | A paragraph of conditioning; few-shot barely fits. |
| 2020 | GPT-3 | 2K | Few-shot prompting works — the pattern-in/pattern-out era. |
| The 8K–32K era — documents fit | |||
| 2022 | GPT-3.5 / early GPT-4 | 4K | Chat with history; whole emails. |
| 2023 | GPT-4 32K · Claude 1 100K | 32K–100K | Whole documents and small codebases — 'just paste it' begins; RAG's first challenger. |
| The 128K–1M era — the corpus fits | |||
| 2023–24 | GPT-4 Turbo | 128K | Book-length context as a standard tier; needle-in-haystack benchmarks go mainstream. |
| 2024 | Gemini 1.5 Pro | 1M | Hour-long video, whole repos — retrieval-optional for many tasks. |
| 2025+ | Frontier models | 200K–1M+ | Long context as table stakes; cache pricing makes giant prompts affordable. |
| The caveats that survived every jump | |||
| — | lost in the middle | — | Recall is U-shaped: start and end of context beat the middle. |
| — | cost & latency | — | Attention cost grows with length; long context bills per token. |
| — | prompt caching | — | Stable prefixes cache at ~10% — architecture matters again (see the Cache Breakpoint Planner). |