Lesson 4 / 25

Evaluator-Optimizer Loops

Improve output through a generate, critique, revise loop with a clear stopping rule.

Generate, judge, revise

One call produces a draft, another judges it against clear criteria, and the feedback goes back for a revision. It works when you can state what "good" means, for example "all tests pass" or "meets the style guide". Without a clear criterion the loop only adds cost.

A bounded loop

Always cap the number of rounds. Here the loop stops when the evaluator approves or after three tries.

def refine(task: str, max_rounds: int = 3) -> str:
    draft = llm(task)
    for _ in range(max_rounds):
        verdict = llm(f"Does this meet the criteria? Reply OK or list problems:\n{draft}")
        if verdict.strip().startswith("OK"):
            break
        draft = llm(f"Revise to fix these problems:\n{verdict}\n\nDraft:\n{draft}")
    return draft

Prefer objective checks

A test suite or a linter is a more reliable evaluator than another model's opinion. Use a model judge only for qualities code cannot measure.

Quick check: What must every evaluator-optimizer loop have?

  • A maximum number of rounds
  • An unlimited retry count
  • No criteria at all
  • A different language per round
Answer

A maximum number of rounds — A cap prevents infinite loops and runaway cost when the criteria are never met.