What it does
How you split a document changes what your retriever finds. The RAG Chunk Comparator chunks one document three ways — fixed-size (the naive default), sentence-aware (boundaries never cut a sentence), and markdown-aware (sections stay together, headings carried as context) — then compares chunk counts, size spread, and the share of boundaries that land on sentence ends. Mid-sentence cuts are the classic recall killer: half a thought embeds badly. Runs 100% client-side.
How to use it
- Paste the document your retriever will ingest.
- Set the target chunk size in tokens (300–500 is a common starting point).
- Tap the three stat tiles to switch strategies and read each one’s chunks — the sentence- boundary percentage tells you which strategies cut thoughts in half.
Examples
A 200-token policy document at a 60-token target:
fixed 4 chunks · 41–60 tok · 33% sentence boundaries
sentence 4 chunks · 38–58 tok · 100% sentence boundaries
markdown 3 chunks · 52–60 tok · 100% sentence boundaries (headings kept)
Fixed wins on uniformity; markdown wins on coherence — every chunk is a complete section with its heading attached.
Good to know
- Sentence and markdown strategies trade exact size for coherence: a single sentence larger than the target becomes its own oversize chunk rather than being split.
- Markdown chunks carry the nearest heading, which doubles as cheap metadata for citation and filtering downstream.
- Token counts are prose-heuristic estimates (~15% tolerance).
- Runs 100% client-side. Your document never leaves the browser.
- Related: Embedding Chunk Planner, Token Estimator, Context Window Planner.