Lesson 20 / 25
Checkpoints and Resume
Save loop state after each step so a stopped or crashed run can continue.
State you can reload
A long run will eventually hit a limit, a deploy or a crash. If you save the messages, the step counter and the budget used after every successful step, a new process can load them and continue, spending only what remains. Make tool actions idempotent (safe to run twice) where you can, because a resume may repeat the last step.
Save and load
Writing to a temporary file and renaming it avoids a half-written checkpoint if the process dies mid-write.
import json, os
def save(path, state):
tmp = path + ".tmp"
with open(tmp, "w") as f:
json.dump(state, f)
os.replace(tmp, path) # atomic on the same filesystem
def load(path):
with open(path) as f:
return json.load(f)Quick check: Why make tool actions idempotent when resuming is possible?
- Resuming may repeat the last step
- Idempotent tools are faster
- The API requires it
- It avoids all errors
Answer
Resuming may repeat the last step — If the last step ran but was not recorded, a second run must not cause duplicate effects.