Document Splitting

Chunking cuts long documents into pieces suitable for retrieval: embedding models have limited input, and the goal of retrieval is to inject only relevant content into context.

Three strategies:

  • fixed size with overlap — simple, but cuts structure;
  • recursive by natural boundaries — the standard choice;
  • semantic — cut where the similarity between adjacent sentences drops.

Typical size 256–1024 tokens, overlap 10–20%.

An inherent downside: the chunk loses connection to the original context — this is treated by context-aware retrieval.

Related: RAG, Context-aware retrieval, Dense embedding