Lesson 25 / 27

Serving, Batching and Choosing a Model

Compare hosted APIs and self-hosted open models, and understand throughput.

Rent or run

Hosted APIs give the strongest models with no infrastructure, pay-per-token pricing and fast iteration, but your data goes to a provider and costs scale with use. Open-weight models (run on your servers or a managed host) give control, privacy and predictable cost at high volume, but you handle GPUs, scaling and updates. Serving engines improve throughput with continuous batching (adding new requests to a running batch), paged KV-cache management and streaming tokens to the user as they are produced. Decide with a small evaluation on your own data: quality, latency, cost, privacy and licence terms. Keep your code behind a thin interface so you can swap models later.

A decision table

Starting points; confirm with your own tests.

Situation                                   Lean toward
Prototype, small volume, need best quality    hosted API
Sensitive data that cannot leave your network open-weight, self-hosted
Very high volume, simple repetitive task      smaller / quantized open model
Need latest knowledge or citations            any model + retrieval
Unsure                                        run an eval on 50 real examples

Quick check: What is a benefit of streaming tokens?

  • Answers become more accurate
  • The model becomes smaller
  • Users see output start sooner, improving perceived latency
  • Tokens become free
Answer

Users see output start sooner, improving perceived latency — Streaming shows partial output immediately instead of waiting for the full answer.