| Strategies |
|---|
| fixed-size | Every N tokens, optionally with overlap. | Uniform sources, simple pipelines, predictable budgets. | Cuts thoughts mid-sentence — the classic recall killer. |
| sentence-aware | Groups whole sentences up to the target size. | Prose without structure — support docs, articles. | One long sentence can blow past the target. |
| markdown / heading | Splits at headings; oversized sections fall back to sentences. | Docs with real structure — READMEs, specs, wikis. | Heading-less or deeply nested documents degrade. |
| semantic | Embeds sentences, splits where embedding similarity drops. | Topic-shifting text where structure is absent. | Slower and tuny; needs a good embedding model up front. |
| late chunking | Embeds the FULL document first, then pools token spans into chunks. | Long docs where context changes word meaning. | Needs long-context embedders and custom pooling. |
| Sizing |
|---|
| 256–512 tokens | The common target band for retrieval chunks. | Default starting point for most corpora. | Not a law — match your embedder's sweet spot. |
| overlap 10–20% | Repeat a tail at the start of the next chunk. | Softens boundary loss on fixed chunking. | Duplicates in the index; overcounts recall hits. |
| small-to-big | Retrieve on small chunks, return the surrounding section. | Precision of small chunks + context of big ones. | Needs parent pointers in storage. |
| Pitfalls that quietly kill recall |
|---|
| mid-sentence cuts | Half a thought embeds badly and cites worse. | — | Check the sentence-boundary share of your splitter. |
| orphaned headings | The heading lands in chunk N, its content in N+1. | — | Carry the heading into every chunk of its section. |
| table shredding | Tables split across chunks lose their header row. | — | Split tables on rows, never mid-table. |
| noise chunks | Boilerplate, nav, and legal footers become top hits. | — | Filter junk before chunking, not after retrieval. |