Lesson 3 / 25

Parallelization and Orchestrator-Workers

Run independent sub-tasks in parallel and let an orchestrator divide work it cannot predict in advance.

Fan out, then combine

Parallelization runs independent calls at the same time, either on different sections of a task or several attempts at the same task to compare. In orchestrator-workers, one model decides at runtime how to split a task, hands pieces to worker calls, and merges their results. Use it when you cannot know the sub-tasks in advance, such as editing several unknown files.

Parallel calls with asyncio

gather starts all reviews together, so total time is about the slowest one rather than the sum. async_llm is any async model call.

import asyncio

async def review_all(files: list[str]) -> list[str]:
    tasks = [async_llm(f"Review this file:\n{f}") for f in files]
    return await asyncio.gather(*tasks)

A head chef

A head chef reads the orders, decides which dishes go to which cook, and plates the final meal. The orchestrator does the same with sub-tasks and worker calls.

Quick check: When is orchestrator-workers a better fit than a fixed chain?

  • When the sub-tasks are known in advance
  • When you want no model calls
  • When the sub-tasks depend on the input and cannot be listed upfront
  • When latency does not matter at all
Answer

When the sub-tasks depend on the input and cannot be listed upfront — The orchestrator discovers the work at runtime, which a fixed chain cannot do.