Lesson 6 / 27
Chunking Strategies
Split text into retrievable pieces with the right size and overlap.
Not too big, not too small
Retrieval works on chunks, so their shape matters. Too small and a chunk loses context ("it", "the above") and may not contain a full answer. Too large and its embedding blurs several topics, retrieval is less precise, and prompts become expensive. Typical starting points are 200 to 500 tokens with 10 to 20% overlap, but the best size depends on your documents and questions, so test it. Structure-aware chunking is usually better than fixed windows: split on headings, paragraphs or sentences, keep a table or code block whole, and prepend the document and section title to each chunk so it makes sense alone.
Fixed windows vs sentence chunks, run
I ran this plain-Python (standard library only) example. Fixed 60-character windows with 15 overlap cut words in half ("retriev", "nks are"), while sentence-based chunking keeps whole sentences together.
text = ("RAG splits documents into chunks. Each chunk is embedded. "
"At question time the nearest chunks are retrieved. "
"They are placed in the prompt. The model answers from them.")
def fixed(text, size, overlap):
step = size - overlap
return [text[i:i + size] for i in range(0, len(text), step) if text[i:i + size].strip()]
def by_sentence(text, max_chars):
out, cur = [], ""
for s in text.split(". "):
s = s.strip().rstrip(".") + "."
if cur and len(cur) + len(s) + 1 > max_chars:
out.append(cur); cur = s
else:
cur = (cur + " " + s).strip()
if cur: out.append(cur)
return out
for c in fixed(text, 60, 15): print(repr(c))
print("---")
for c in by_sentence(text, 70): print(repr(c))
Output:
'RAG splits documents into chunks. Each chunk is embedded. At' 'is embedded. At question time the nearest chunks are retriev' 'nks are retrieved. They are placed in the prompt. The model ' 'mpt. The model answers from them.' --- 'RAG splits documents into chunks. Each chunk is embedded.' 'At question time the nearest chunks are retrieved.' 'They are placed in the prompt. The model answers from them.'
Test chunk size on real questions
Try 2 or 3 sizes and compare recall on a fixed set of questions (see the evaluation section). Do not guess.
Quick check: What is a problem with chunks that are too large?
- They remove metadata
- They cannot be stored
- They always improve accuracy
- Their embeddings blur several topics and prompts get costly
Answer
Their embeddings blur several topics and prompts get costly — Large chunks reduce retrieval precision and waste prompt tokens.