All posts
// / Blog

The most impactful change I made to a RAG system wasn't the model, the embeddings, or the…

retrieval algorithm. It was how I chunked the documents.

Default chunking (500 tokens with 50 overlap) is the "Hello World" of RAG. It works for demos. It fails in production.

What I've learned through painful iteration:

Semantic chunking — split at paragraph or section boundaries, not arbitrary token counts. Documents have natural structure. Respect it.

Context-enriched chunks — prepend the document title and section header to every chunk. "Revenue was $5.2M" means nothing without knowing it's from "Q3 2025 Financial Report > North America."

Hierarchical chunking — small chunks for precise retrieval, parent chunks for context. Retrieve the small chunk, but pass the parent chunk to the LLM. Best of both worlds.

Table-aware chunking — tables need special handling. A table split across two chunks is useless. Detect tables, keep them intact, and add a natural language summary.

The right chunking strategy depends entirely on your documents. Financial reports need different chunking than legal contracts, which need different chunking than technical documentation.

There's no universal best practice. Only informed experimentation. But getting chunking right typically improves RAG quality more than swapping the LLM.

#RAG#Chunking#LLM#AIArchitecture#NLP#MachineLearning