# Serving, Batching and Choosing a Model — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/d-serving

> Compare hosted APIs and self-hosted open models, and understand throughput.

## Rent or run

**Hosted APIs** give the strongest models with no infrastructure, pay-per-token pricing and fast iteration, but your data goes to a provider and costs scale with use. **Open-weight models** (run on your servers or a managed host) give control, privacy and predictable cost at high volume, but you handle GPUs, scaling and updates. Serving engines improve throughput with **continuous batching** (adding new requests to a running batch), **paged KV-cache** management and **streaming** tokens to the user as they are produced. Decide with a small evaluation on your own data: quality, latency, cost, privacy and licence terms. Keep your code behind a thin interface so you can swap models later.

## A decision table

Starting points; confirm with your own tests.

```text
Situation                                   Lean toward
Prototype, small volume, need best quality    hosted API
Sensitive data that cannot leave your network open-weight, self-hosted
Very high volume, simple repetitive task      smaller / quantized open model
Need latest knowledge or citations            any model + retrieval
Unsure                                        run an eval on 50 real examples
```

**Quiz:** What is a benefit of streaming tokens?

- [ ] Answers become more accurate
- [ ] The model becomes smaller
- [x] Users see output start sooner, improving perceived latency
- [ ] Tokens become free

*Answer:* Users see output start sooner, improving perceived latency. Streaming shows partial output immediately instead of waiting for the full answer.
