Lesson 9 / 27
How Streaming Works
Understand server-sent events and consume them with each SDK.
Events instead of one big response
A long answer can take many seconds to generate. With streaming the server sends the reply as a series of small events (using Server-Sent Events over a normal HTTP response) while the model is still generating, so a user sees the first words almost at once. Total time is similar, but perceived latency drops sharply. Anthropic's stream has typed events (message start, content-block start, text deltas, content-block stop, message delta with the stop reason and usage, message stop); OpenAI streams chunks whose delta holds new text, ending with a final chunk and [DONE]. SDKs hide the parsing: you iterate over text pieces and can ask for the final assembled message and usage at the end.
Show text as it is produced
Streaming sends the reply in small events so users see it immediately.
Streaming with both SDKs, run
I ran this in a Python virtual environment with the anthropic 1.11.0 and openai 3.22.1 SDKs against a small local stand-in server (shown in the testing topic, saved as mock.py). The server returns canned replies, so no key, network or real model is involved: it proves how the SDK builds requests and handles replies, not what a real model would say. Both SDKs deliver the same three text pieces ("Hel", "lo ", "there"). Anthropic also provides the final assembled message with its stop reason; with OpenAI you join the delta.content pieces yourself.
import mock, anthropic, openai
srv = mock.start(); base = f"http://127.0.0.1:{srv.server_address[1]}" # local stand-in server, not a real API
a = anthropic.Anthropic(api_key="k", base_url=base)
with a.messages.stream(model="demo-model", max_tokens=50, messages=[{"role": "user", "content": "Hi"}]) as stream:
pieces = list(stream.text_stream)
final = stream.get_final_message()
print("anthropic pieces:", pieces, "| final:", final.content[0].text, "|", final.stop_reason)
o = openai.OpenAI(api_key="k", base_url=base + "/v1")
chunks = o.chat.completions.create(model="demo-model", stream=True, messages=[{"role": "user", "content": "Hi"}])
parts = []
for ch in chunks:
delta = ch.choices[0].delta.content
if delta: parts.append(delta)
print("openai pieces: ", parts, "| joined:", "".join(parts))
Output:
anthropic pieces: ['Hel', 'lo ', 'there'] | final: Hello there | end_turn openai pieces: ['Hel', 'lo ', 'there'] | joined: Hello there
Handle the end of the stream
Read the final event for the stop reason and usage; a stream that ends without it may have been cut short.
Quick check: What does streaming mainly improve?
- The token price
- The accuracy of answers
- Perceived latency, because text appears while it is generated
- The key security
Answer
Perceived latency, because text appears while it is generated — Users see the first words sooner even if total time is similar.