# Hosted APIs vs Self-Hosting — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/s-host

> Compare control, cost, effort and risk.

## Rent first, own when the numbers say so

**Hosted APIs** give the strongest models, no infrastructure to run, fast iteration and pay-per-use pricing; the trade-offs are per-token cost at high volume, data leaving your network, dependence on the provider's uptime, rate limits and model deprecations. **Self-hosting an open-weight model** gives control over data, versions, customisation and (at high steady volume) unit cost, but you take on GPUs, serving software, scaling, upgrades, monitoring and on-call, and the strongest models may not be open. Many teams start hosted, measure real volume and quality needs, and move selected workloads to smaller self-hosted or fine-tuned models when the economics, privacy or latency justify it. Whatever you choose, keep a **thin abstraction** over the model call so switching is a configuration change, and keep an **evaluation set** that lets you compare candidates fairly.

## Hosted or self-hosted, gradual and reversible

Choose how to run the model, understand what limits throughput, and release changes gradually with a way back.

![Three decisions: where it runs, what limits it, how it ships.](assets/figures/llm-engineering/section-6-map.svg) — Figure 6.1 — Where it runs, what limits it and how it ships.

## Hosted versus self-hosted at a glance

A starting comparison; your own numbers decide.

```text
                      hosted API                          self-hosted open model
quality ceiling       strongest available models           good, depends on the open model chosen
effort to start       hours                                 days to weeks (GPUs, serving, scaling)
unit cost             pay per token; high at huge volume   fixed hardware; cheap at high, steady volume
data control          leaves your network (check terms)    stays in your environment
customisation         prompts, some fine-tuning options     full control: weights, adapters, quantisation
risks                 outages, rate limits, deprecations   you own ops: scaling, security, upgrades, on-call
```

## Measure real volume first

Decide on self-hosting from measured traffic and costs, not guesses.

**Quiz:** What helps most when you might switch between models or providers later?

- [ ] Never reading release notes
- [ ] Hard-coding one provider's details everywhere
- [ ] Avoiding tests
- [x] A thin abstraction over the model call and an evaluation set

*Answer:* A thin abstraction over the model call and an evaluation set. Abstraction makes switching cheap; evaluation makes it safe.
