Lesson 8 / 28
Preparing Text: Chunking, Truncation, Query vs Document
Avoid silent quality losses before embedding.
Garbage in, blurred vectors out
Embedding quality depends heavily on what you feed in. Chunk long documents into coherent pieces (a few hundred tokens, split on headings or paragraphs, with a little overlap): one vector per giant document blurs many topics. Respect the model's maximum input length, because longer text is usually truncated silently. Clean boilerplate, navigation text and encoding errors. Prepend titles or section names to chunks so they make sense alone. Some models are asymmetric and expect different prefixes or instructions for queries and documents (for example "query: ..." and "passage: ..."): follow the model card or retrieval quality drops. Batch requests for throughput, normalise vectors if required, and cache embeddings so unchanged text is never embedded twice.
Query and passage prefixes (illustrative)
Some retrieval models expect role markers; the exact strings are model-specific, so use those in the model card. Not run here.
docs = ["passage: Employees get 24 days of annual leave."] # document side
query = "query: how many vacation days do I get?" # query side
# doc_vectors = model.encode(docs, normalize_embeddings=True, batch_size=64)
# q_vector = model.encode([query], normalize_embeddings=True)
# scores = doc_vectors @ q_vector.T # normalised -> dot == cosineQuick check: What often happens to text longer than a model's maximum input length?
- It is translated
- It is embedded perfectly
- It is silently truncated, losing the end
- The model trains on it
Answer
It is silently truncated, losing the end — Chunk below the limit so nothing important is cut.