Lesson 1 / 27
What an LLM API Is
See the API as a stateless web service that turns messages into a reply.
A web service with one main call
A hosted language model is reached through a web API: you send an HTTPS POST with a JSON body (model name, messages, settings) and an API key, and receive JSON containing the model's reply and usage (how many input and output tokens were used). The service is stateless: it remembers nothing between calls, so your application sends the whole conversation each time. Billing is per token, so every request has a price and a latency. The two most widely used families are Anthropic's Claude API (the Messages API) and OpenAI's API (Chat Completions and the newer Responses API). Their details differ, but the ideas in this course apply to both and to most other providers.
Request in, message out
An LLM API is an HTTPS call: you send messages and settings, and get a message and usage numbers back.
A very forgetful helpdesk
Each call is a fresh phone call to a helpdesk agent who remembers nothing. You must repeat the relevant history every time you call.
The shape of a call
Both providers follow this pattern, with different field names.
POST https://<provider-host>/<messages-endpoint>
headers: API key (+ content-type, + version header for some providers)
body: { "model": "<model-name>",
"messages": [ {"role": "user", "content": "..."} ],
"max_tokens": ..., "temperature": ..., ... }
response: { "content"/"choices": [ the reply ],
"stop_reason"/"finish_reason": "...",
"usage": { input_tokens, output_tokens } }Quick check: What does "stateless" mean for an LLM API?
- It only works offline
- It stores your conversation forever for free
- It remembers nothing between calls, so you resend the conversation
- It has no billing
Answer
It remembers nothing between calls, so you resend the conversation — Each request must carry all the context the model needs.