# Preparing Text: Chunking, Truncation, Query vs Document — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/c-prep

> Avoid silent quality losses before embedding.

## Garbage in, blurred vectors out

Embedding quality depends heavily on what you feed in. **Chunk** long documents into coherent pieces (a few hundred tokens, split on headings or paragraphs, with a little overlap): one vector per giant document blurs many topics. Respect the model's **maximum input length**, because longer text is usually **truncated silently**. **Clean** boilerplate, navigation text and encoding errors. Prepend titles or section names to chunks so they make sense alone. Some models are **asymmetric** and expect different prefixes or instructions for **queries** and **documents** (for example "query: ..." and "passage: ..."): follow the model card or retrieval quality drops. **Batch** requests for throughput, **normalise** vectors if required, and **cache** embeddings so unchanged text is never embedded twice.

## Query and passage prefixes (illustrative)

Some retrieval models expect role markers; the exact strings are model-specific, so use those in the model card. Not run here.

```python
docs = ["passage: Employees get 24 days of annual leave."]          # document side
query = "query: how many vacation days do I get?"                 # query side

# doc_vectors = model.encode(docs, normalize_embeddings=True, batch_size=64)
# q_vector    = model.encode([query], normalize_embeddings=True)
# scores      = doc_vectors @ q_vector.T          # normalised -> dot == cosine
```

**Quiz:** What often happens to text longer than a model's maximum input length?

- [ ] It is translated
- [ ] It is embedded perfectly
- [x] It is silently truncated, losing the end
- [ ] The model trains on it

*Answer:* It is silently truncated, losing the end. Chunk below the limit so nothing important is cut.
