Lesson 20 / 25

Checkpoints and Resume

Save loop state after each step so a stopped or crashed run can continue.

State you can reload

A long run will eventually hit a limit, a deploy or a crash. If you save the messages, the step counter and the budget used after every successful step, a new process can load them and continue, spending only what remains. Make tool actions idempotent (safe to run twice) where you can, because a resume may repeat the last step.

Save and load

Writing to a temporary file and renaming it avoids a half-written checkpoint if the process dies mid-write.

import json, os

def save(path, state):
    tmp = path + ".tmp"
    with open(tmp, "w") as f:
        json.dump(state, f)
    os.replace(tmp, path)          # atomic on the same filesystem

def load(path):
    with open(path) as f:
        return json.load(f)

Quick check: Why make tool actions idempotent when resuming is possible?

  • Resuming may repeat the last step
  • Idempotent tools are faster
  • The API requires it
  • It avoids all errors
Answer

Resuming may repeat the last step — If the last step ran but was not recorded, a second run must not cause duplicate effects.