# Allocating the Context Window — Agent Loops, Stop Conditions and Token Budgets

Source: https://www.geekswithgeeks.com/en/agent-loops/budget-allocation

> Split the window among system prompt, tools, history and an output reserve.

## Plan the shares

Think of the window as a fixed pie. The system prompt and tool definitions are a **fixed cost** paid every turn. History is the **growing share**. The reply needs a **reserve**, because if history fills the window, there is no room left to answer. Decide a trigger, for example "start trimming when history passes 60% of the window".

## A planning table

The numbers are an example for a 200,000-token window. Adjust to your model and task.

```text
Window                       200,000
- System prompt + tool defs    -6,000   (fixed, every turn)
- Output reserve              -16,000   (max_tokens + safety)
= Available for history       178,000
Start trimming at 60%         ~107,000
```

## Tool definitions add up

Fifty tools with long descriptions can cost thousands of tokens every turn. Load only the tools relevant to the current task.

**Quiz:** Why reserve space for the model's reply?

- [ ] Replies are free
- [x] If history fills the window there is no room left to answer
- [ ] It makes tools faster
- [ ] The API requires a reserve flag

*Answer:* If history fills the window there is no room left to answer. Input and output share the window, so the reply needs guaranteed room.
