Lesson 9 / 31
Streaming and Async
Show output as it is produced and run calls concurrently.
Faster to first word
Language models generate token by token, so waiting for the whole answer feels slow. Every runnable has stream (and async astream) that yields pieces as they arrive, letting a UI show text immediately; perceived latency drops even though total time is similar. The async methods (ainvoke, abatch, astream) let a server handle many requests concurrently without blocking threads. When streaming through a parser, make sure the parser supports partial output (StrOutputParser does; a JSON parser may emit partial objects). Remember to handle errors that occur mid-stream.
Streaming pieces, run
I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. The fake model streams its canned reply one character at a time (15 pieces here); real models stream tokens. Joining the pieces gives the full text.
from langchain_core.language_models.fake_chat_models import FakeListChatModel
from langchain_core.output_parsers import StrOutputParser
chain = FakeListChatModel(responses=["streaming works"]) | StrOutputParser()
pieces = list(chain.stream("hi"))
print(pieces[:5], "...", len(pieces), "pieces")
print("".join(pieces))
Output:
['s', 't', 'r', 'e', 'a'] ... 15 pieces streaming works
Quick check: What does streaming improve?
- The vocabulary size
- The model's accuracy
- Perceived latency: users see text sooner
- The training data
Answer
Perceived latency: users see text sooner — Streaming shows partial output immediately.