Lesson 22 / 29
Hosted APIs vs Self-Hosting
Compare control, cost, effort and risk.
Rent first, own when the numbers say so
Hosted APIs give the strongest models, no infrastructure to run, fast iteration and pay-per-use pricing; the trade-offs are per-token cost at high volume, data leaving your network, dependence on the provider's uptime, rate limits and model deprecations. Self-hosting an open-weight model gives control over data, versions, customisation and (at high steady volume) unit cost, but you take on GPUs, serving software, scaling, upgrades, monitoring and on-call, and the strongest models may not be open. Many teams start hosted, measure real volume and quality needs, and move selected workloads to smaller self-hosted or fine-tuned models when the economics, privacy or latency justify it. Whatever you choose, keep a thin abstraction over the model call so switching is a configuration change, and keep an evaluation set that lets you compare candidates fairly.
Hosted or self-hosted, gradual and reversible
Choose how to run the model, understand what limits throughput, and release changes gradually with a way back.
Hosted versus self-hosted at a glance
A starting comparison; your own numbers decide.
hosted API self-hosted open model
quality ceiling strongest available models good, depends on the open model chosen
effort to start hours days to weeks (GPUs, serving, scaling)
unit cost pay per token; high at huge volume fixed hardware; cheap at high, steady volume
data control leaves your network (check terms) stays in your environment
customisation prompts, some fine-tuning options full control: weights, adapters, quantisation
risks outages, rate limits, deprecations you own ops: scaling, security, upgrades, on-callMeasure real volume first
Decide on self-hosting from measured traffic and costs, not guesses.
Quick check: What helps most when you might switch between models or providers later?
- Never reading release notes
- Hard-coding one provider's details everywhere
- Avoiding tests
- A thin abstraction over the model call and an evaluation set
Answer
A thin abstraction over the model call and an evaluation set — Abstraction makes switching cheap; evaluation makes it safe.