# Failure Handling and Stop Conditions — Advanced Agent Workflows and Skills

Source: https://www.geekswithgeeks.com/en/agent-workflows/rel-failure-limits

> Add retries, timeouts, step limits and human checkpoints so agents fail safely.

## Plan for things going wrong

Agents loop, retry the same failing step, or wander. Defend with a **maximum number of steps**, a **time limit**, a **cost cap**, and **checkpoints** where a human approves risky actions such as deleting data or deploying. When a limit is hit, stop and report clearly instead of continuing.

## A guarded agent loop

`step` runs one model-and-tool turn and `is_done` checks the result. The loop cannot run forever or spend without limit.

```python
MAX_STEPS, MAX_COST = 15, 2.00   # dollars
spent = 0.0
for i in range(MAX_STEPS):
    result, cost = step(state)
    spent += cost
    if is_done(result):
        break
    if spent > MAX_COST:
        raise RuntimeError("Cost cap reached; stopping for review")
else:
    raise RuntimeError("Step limit reached without finishing")
```

## Make failures visible

Log every tool call and decision. When something goes wrong at 3 a.m., a readable trace is the difference between a ten-minute fix and a day of guessing.

**Quiz:** Which is a good stop condition for an autonomous loop?

- [ ] Run until the model says it is finished, no matter what
- [x] A maximum number of steps plus a cost cap
- [ ] Never stop
- [ ] Stop at random

*Answer:* A maximum number of steps plus a cost cap. Hard limits guarantee the loop ends even when the model never decides it is done.
