Lesson 4 / 25
Evaluator-Optimizer Loops
Improve output through a generate, critique, revise loop with a clear stopping rule.
Generate, judge, revise
One call produces a draft, another judges it against clear criteria, and the feedback goes back for a revision. It works when you can state what "good" means, for example "all tests pass" or "meets the style guide". Without a clear criterion the loop only adds cost.
A bounded loop
Always cap the number of rounds. Here the loop stops when the evaluator approves or after three tries.
def refine(task: str, max_rounds: int = 3) -> str:
draft = llm(task)
for _ in range(max_rounds):
verdict = llm(f"Does this meet the criteria? Reply OK or list problems:\n{draft}")
if verdict.strip().startswith("OK"):
break
draft = llm(f"Revise to fix these problems:\n{verdict}\n\nDraft:\n{draft}")
return draftPrefer objective checks
A test suite or a linter is a more reliable evaluator than another model's opinion. Use a model judge only for qualities code cannot measure.
Quick check: What must every evaluator-optimizer loop have?
- A maximum number of rounds
- An unlimited retry count
- No criteria at all
- A different language per round
Answer
A maximum number of rounds — A cap prevents infinite loops and runaway cost when the criteria are never met.