Playbooks

A chunking strategy you can defend in review

Chunk size is usually set by a library default and defended by nothing. Choose it from the documents themselves, and be able to explain the choice to someone who will ask why.

Start from the failure you are preventing

Chunking has exactly two failure directions. Chunks that are too small lose the context needed to interpret them, so the retrieval step finds the right passage and the model misreads it. Chunks that are too large dilute the embedding, so the retrieval step misses the passage because a paragraph of irrelevant text dominates the vector.

Which direction hurts more depends on the document type, and that is the whole design decision. Narrative documents tolerate larger chunks because meaning is distributed. Reference documents — policies, contracts, specification tables — need smaller ones because each clause is self-contained and must be retrieved alone.

Chunk on structure before you chunk on length

Length-based splitting is a fallback, not a strategy. Documents already have structure: headings, sections, numbered clauses, table rows. Splitting on that structure gives chunks that correspond to units a human would cite, and citations are what the answer needs to sound credible.

  • Split at headings and clause boundaries first. Use length as a constraint on top of that, with an overlap large enough to carry a sentence across the boundary.
  • Keep tables whole when they are small, and split them by row with the header repeated when they are not. A value separated from its column header is worse than a missing value.
  • Never split a rule from its exception. If the exception is a following sentence or a sub-clause, they belong in the same chunk even if that makes it longer than target.
  • Prepend the section path to every chunk — document title, section, subsection. A chunk is retrieved out of context and must be readable in isolation, in the model's context window, without the original around it.

Choosing a size, not a default

Take a representative sample of real documents, split them with two or three candidate configurations, and inspect the resulting chunks by reading them. The question to answer while reading is concrete: could a competent person answer a question from this chunk alone? If the answer is no for a meaningful share of chunks, the configuration is wrong regardless of what the evaluation score says.

Then check the interaction with retrieval. Short chunks improve precision on specific questions and hurt questions that require surrounding context. If the corpus contains both kinds of question, that is the honest argument for structural chunking with variable size rather than one global number.

What to write down so the choice survives

  • The configuration, the reason for it, and the document families it was chosen for.
  • The two or three questions that motivated the choice, kept as evaluation items so a future change is measured against them.
  • What to do for a new document family, since the configuration is a decision about content, not a global constant.
  • The date and the corpus it was tuned on, because chunking configured against one corpus is not automatically right for another.

What to do about it

  • Chunk on document structure first, length second
  • Attach context to every chunk so it can be read alone
  • Never split a rule from its exception or a value from its header