Lesson 13 / 25

Clipping Tool Results

Cut long tool outputs to a safe size while telling the model what was removed.

Big outputs are the usual culprit

One careless tool result, like a 50,000-line log file, can use more tokens than the rest of the conversation combined, and it is then resent on every later turn. Clip results to a limit and say so, so the model knows to ask for a narrower view.

Keep only what matters

Trim, summarise and offload so the history stays useful without growing forever.

Three tools: clip results, summarise old turns, store externally.
Figure 4.1 — Clip, summarise and offload.

A clip helper

This ran as shown: the output ends with a note saying how many characters were removed.

def clip(text, limit=200):
    if len(text) <= limit:
        return text
    return text[:limit] + f"\n...[truncated {len(text) - limit} chars]"

print(clip("x" * 250, 200)[-30:])

Output:

xxxxxx
...[truncated 50 chars]

Make tools paginate

Better than clipping is a tool that returns the first page and a way to ask for more, such as offset and limit parameters, so the model pulls only what it needs.

Quick check: Why is one huge tool result so costly?

  • It is resent on every later turn
  • Large results are encrypted
  • It slows the keyboard
  • It is billed once only
Answer

It is resent on every later turn — History is resent each turn, so a big result keeps costing tokens until removed.