Lesson 25 / 27
Serving, Batching and Choosing a Model
Compare hosted APIs and self-hosted open models, and understand throughput.
Rent or run
Hosted APIs give the strongest models with no infrastructure, pay-per-token pricing and fast iteration, but your data goes to a provider and costs scale with use. Open-weight models (run on your servers or a managed host) give control, privacy and predictable cost at high volume, but you handle GPUs, scaling and updates. Serving engines improve throughput with continuous batching (adding new requests to a running batch), paged KV-cache management and streaming tokens to the user as they are produced. Decide with a small evaluation on your own data: quality, latency, cost, privacy and licence terms. Keep your code behind a thin interface so you can swap models later.
A decision table
Starting points; confirm with your own tests.
Situation Lean toward
Prototype, small volume, need best quality hosted API
Sensitive data that cannot leave your network open-weight, self-hosted
Very high volume, simple repetitive task smaller / quantized open model
Need latest knowledge or citations any model + retrieval
Unsure run an eval on 50 real examplesQuick check: What is a benefit of streaming tokens?
- Answers become more accurate
- The model becomes smaller
- Users see output start sooner, improving perceived latency
- Tokens become free
Answer
Users see output start sooner, improving perceived latency — Streaming shows partial output immediately instead of waiting for the full answer.