Lesson 9 / 29

The Context Window as a Budget

Select the most relevant material that fits.

Everything competes for the same space

The context window is the maximum amount of text (measured in tokens) the model can consider in one call. It must hold the instructions, the tool descriptions, the conversation so far (including every tool result), the files the agent has read, and room for the reply. Long contexts are slower and costlier, and models can use information buried in a very long context less reliably, so more is not always better. The harness therefore selects context: it ranks files by relevance to the task (names, symbols, search hits, recent edits), includes the top ones within a token budget, and shows only relevant line ranges instead of whole files. You can help by naming the files that matter, pointing to similar existing code, and keeping tasks focused on one change.

A small window, chosen well

The agent only knows what is in its context, so what you put there, and how it is kept tidy, decides quality.

Four tools: pack, instruct, compact, delegate.
Figure 3.1 — Pack, instruct, compact and delegate.

Packing relevant files into a token budget, run

I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. Files are scored by how often the query terms appear (file names count triple). The pricing file, its tests and the cart fit in 150 of 300 tokens; the long unrelated history document scores zero and is skipped. Real harnesses use richer signals, but the idea is the same.

def estimate_tokens(text): return max(1, len(text) // 4)          # rough: 4 characters per token

def pack_context(files, query_terms, budget):
    """Pick the most relevant files that fit the token budget."""
    def score(item):
        name, text = item
        return sum(text.lower().count(t) + 3 * name.lower().count(t) for t in query_terms)
    chosen, used = [], 0
    for name, text in sorted(files.items(), key=lambda kv: -score(kv)):
        cost = estimate_tokens(text)
        if score((name, text)) == 0 or used + cost > budget:
            continue
        chosen.append((name, cost)); used += cost
    return chosen, used

files = {
    "shop/pricing.py": "def apply_discount(price, percent): ... discount discount tax " * 5,
    "shop/cart.py": "class Cart: total discount " * 4,
    "docs/history.md": "release notes and history " * 200,
    "tests/test_pricing.py": "def test_discount ... discount " * 6,
}
chosen, used = pack_context(files, ["discount", "pricing"], budget=300)
print("chosen:", chosen)
print("tokens used:", used, "of 300")

Output:

chosen: [('shop/pricing.py', 77), ('tests/test_pricing.py', 46), ('shop/cart.py', 27)]
tokens used: 150 of 300

Name the files that matter

Telling the agent which files or functions are involved saves search steps and context.

Quick check: Why not always put the whole repository in the context?

  • Because context is free
  • Repositories cannot be read
  • It rarely fits, costs more, is slower and buries the relevant parts
  • Because models dislike code
Answer

It rarely fits, costs more, is slower and buries the relevant parts — Selecting relevant material beats sending everything.