# Context Windows and Token Budgets — Prompt Engineering

Source: https://www.geekswithgeeks.com/en/prompt-engineering/x-budget

> Fit instructions, history and documents in the limit.

## A fixed-size backpack

The **context window** is the maximum number of tokens (prompt plus reply) the model handles in one call. Split the budget deliberately: system instructions, retrieved documents, conversation history, the user question, and **room for the answer**. When something must go, drop or summarise the oldest and least relevant content first and keep the system rules and the latest turns. Always measure with the provider's tokenizer; the 4-characters-per-token rule is only a rough estimate for English, and Hindi often needs more tokens per word.

## What goes in the window

The context window is a budget: choose what to include, in what order.

![Four decisions: budget, order, retrieve, summarise.](assets/figures/prompt-engineering/section-5-map.svg) — Figure 5.1 — Budget, order, retrieve and summarise.

## Trimming history to a budget, run

I ran this plain-Python (standard library only) example. With a 120-token budget, the system prompt and question use some tokens, then the newest messages are added until the budget is full: 4 of 8 messages (the newest, numbers 5 to 8) are kept, using 101 tokens by the rough estimate.

```python
def rough_tokens(text): return max(1, round(len(text) / 4))

def fit(system, history, question, budget):
    used = rough_tokens(system) + rough_tokens(question)
    kept = []
    for msg in reversed(history):          # keep the newest messages first
        t = rough_tokens(msg)
        if used + t > budget: break
        kept.append(msg); used += t
    return list(reversed(kept)), used

history = ["message " + str(i) + " " + "x" * 80 for i in range(1, 9)]
kept, used = fit("You are a helpful assistant.", history, "And what about refunds?", 120)
print("kept", len(kept), "of", len(history), "messages; tokens used:", used)
print([m.split()[1] for m in kept])

```

Output:

```
kept 4 of 8 messages; tokens used: 101
['5', '6', '7', '8']
```

**Quiz:** When history is too long, what is usually dropped first?

- [ ] The user's latest question
- [ ] The system rules
- [x] The oldest, least relevant messages
- [ ] The reserved room for the answer

*Answer:* The oldest, least relevant messages. Old chatter matters least; rules and the current turn matter most.
