Lesson 15 / 27

Choosing Top-k and Managing Context Cost

Balance recall against prompt size, cost and distraction.

More chunks is not always better

A larger k raises the chance that the right chunk is included but increases prompt tokens (cost and latency) and adds distractors that can confuse the model; evidence buried in the middle of long contexts is also used less reliably. Pick k by measuring answer quality at several values, set a token budget for context, apply a minimum relevance threshold so weak matches are dropped (and answer "not found" when nothing passes), and order the chunks sensibly (best first, or by document order). Compress or summarise chunks only when needed, since it can lose detail.

Prompt size by k, run

I ran this plain-Python (standard library only) example. With 300-token chunks, a 150-token system prompt and a 40-token question, k=3 costs 1,090 prompt tokens, k=5 costs 1,690 and k=10 costs 3,190.

chunk_tokens, k, question_tokens, system_tokens, answer = 300, 5, 40, 150, 250
prompt = system_tokens + question_tokens + k * chunk_tokens
print("prompt tokens:", prompt)
for k in (3, 5, 10):
    print("k =", k, "->", system_tokens + question_tokens + k * chunk_tokens, "prompt tokens")

Output:

prompt tokens: 1690
k = 3 -> 1090 prompt tokens
k = 5 -> 1690 prompt tokens
k = 10 -> 3190 prompt tokens

Quick check: Why apply a minimum relevance threshold?

  • So weak matches are not sent, and "not found" can be returned
  • To increase k automatically
  • To disable citations
  • To shorten the index
Answer

So weak matches are not sent, and "not found" can be returned — Irrelevant context invites confident wrong answers.