# Choosing Top-k and Managing Context Cost — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/q-topk

> Balance recall against prompt size, cost and distraction.

## More chunks is not always better

A larger **k** raises the chance that the right chunk is included but increases prompt tokens (cost and latency) and adds **distractors** that can confuse the model; evidence buried in the middle of long contexts is also used less reliably. Pick k by measuring answer quality at several values, set a **token budget** for context, apply a **minimum relevance threshold** so weak matches are dropped (and answer "not found" when nothing passes), and **order** the chunks sensibly (best first, or by document order). Compress or summarise chunks only when needed, since it can lose detail.

## Prompt size by k, run

I ran this plain-Python (standard library only) example. With 300-token chunks, a 150-token system prompt and a 40-token question, k=3 costs 1,090 prompt tokens, k=5 costs 1,690 and k=10 costs 3,190.

```python
chunk_tokens, k, question_tokens, system_tokens, answer = 300, 5, 40, 150, 250
prompt = system_tokens + question_tokens + k * chunk_tokens
print("prompt tokens:", prompt)
for k in (3, 5, 10):
    print("k =", k, "->", system_tokens + question_tokens + k * chunk_tokens, "prompt tokens")

```

Output:

```
prompt tokens: 1690
k = 3 -> 1090 prompt tokens
k = 5 -> 1690 prompt tokens
k = 10 -> 3190 prompt tokens
```

**Quiz:** Why apply a minimum relevance threshold?

- [x] So weak matches are not sent, and "not found" can be returned
- [ ] To increase k automatically
- [ ] To disable citations
- [ ] To shorten the index

*Answer:* So weak matches are not sent, and "not found" can be returned. Irrelevant context invites confident wrong answers.
