# How Streaming Works — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/s-how

> Understand server-sent events and consume them with each SDK.

## Events instead of one big response

A long answer can take many seconds to generate. With **streaming** the server sends the reply as a series of small **events** (using **Server-Sent Events** over a normal HTTP response) while the model is still generating, so a user sees the first words almost at once. Total time is similar, but **perceived latency** drops sharply. Anthropic's stream has typed events (message start, content-block start, text deltas, content-block stop, message delta with the stop reason and usage, message stop); OpenAI streams chunks whose `delta` holds new text, ending with a final chunk and `[DONE]`. SDKs hide the parsing: you iterate over text pieces and can ask for the **final assembled message** and usage at the end.

## Show text as it is produced

Streaming sends the reply in small events so users see it immediately.

![Three ideas: events, assemble, handle errors.](assets/figures/llm-apis/section-3-map.svg) — Figure 3.1 — Events, assemble and handle errors.

## Streaming with both SDKs, run

I ran this in a Python virtual environment with the anthropic 1.11.0 and openai 3.22.1 SDKs against a small local stand-in server (shown in the testing topic, saved as `mock.py`). The server returns canned replies, so no key, network or real model is involved: it proves how the SDK builds requests and handles replies, not what a real model would say. Both SDKs deliver the same three text pieces ("Hel", "lo ", "there"). Anthropic also provides the final assembled message with its stop reason; with OpenAI you join the `delta.content` pieces yourself.

```python
import mock, anthropic, openai
srv = mock.start(); base = f"http://127.0.0.1:{srv.server_address[1]}"      # local stand-in server, not a real API

a = anthropic.Anthropic(api_key="k", base_url=base)
with a.messages.stream(model="demo-model", max_tokens=50, messages=[{"role": "user", "content": "Hi"}]) as stream:
    pieces = list(stream.text_stream)
    final = stream.get_final_message()
print("anthropic pieces:", pieces, "| final:", final.content[0].text, "|", final.stop_reason)

o = openai.OpenAI(api_key="k", base_url=base + "/v1")
chunks = o.chat.completions.create(model="demo-model", stream=True, messages=[{"role": "user", "content": "Hi"}])
parts = []
for ch in chunks:
    delta = ch.choices[0].delta.content
    if delta: parts.append(delta)
print("openai pieces:   ", parts, "| joined:", "".join(parts))

```

Output:

```
anthropic pieces: ['Hel', 'lo ', 'there'] | final: Hello there | end_turn
openai pieces:    ['Hel', 'lo ', 'there'] | joined: Hello there
```

## Handle the end of the stream

Read the final event for the stop reason and usage; a stream that ends without it may have been cut short.

**Quiz:** What does streaming mainly improve?

- [ ] The token price
- [ ] The accuracy of answers
- [x] Perceived latency, because text appears while it is generated
- [ ] The key security

*Answer:* Perceived latency, because text appears while it is generated. Users see the first words sooner even if total time is similar.
