Lesson 26 / 31

Unit Testing with Fake Models

Test your chain logic without calling a real LLM.

Deterministic tests for deterministic code

Most of an LLM app is ordinary code: formatting, routing, parsing, validation, tool functions, retrieval wiring. Test it normally with fake models (FakeListChatModel, MockLLM) and fake embeddings that return fixed outputs. This makes tests fast, free and repeatable, and lets you simulate failures (a malformed JSON reply, a timeout, a tool error). Reserve a separate, smaller set of evaluation runs against the real model for quality, since real outputs are variable. Do not assert on exact free-text answers from a real model in unit tests.

Test without the model, measure with it

Fakes make unit tests fast and free; evaluation sets measure real quality.

Four habits: fake, evaluate, trace, pin.
Figure 7.1 — Fake, evaluate, trace and pin.

A pytest-style test (illustrative)

The chain is built from a prompt, a fake model and a parser, so the test needs no network. The structure matches the chain demonstrated earlier; not run as a test here.

from langchain_core.language_models.fake_chat_models import FakeListChatModel
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate

def build_chain(model):
    prompt = ChatPromptTemplate.from_template("Translate to French: {text}")
    return prompt | model | StrOutputParser()

def test_chain_returns_model_text():
    chain = build_chain(FakeListChatModel(responses=["Bonjour"]))
    assert chain.invoke({"text": "Hello"}) == "Bonjour"

def test_missing_variable_fails_early():
    chain = build_chain(FakeListChatModel(responses=["x"]))
    try:
        chain.invoke({})
        assert False, "expected an error"
    except Exception:
        pass

Quick check: Why use fake models in unit tests?

  • Tests become fast, free and repeatable
  • They improve model quality
  • They remove the need for code
  • They train the embeddings
Answer

Tests become fast, free and repeatable — Fixed outputs make assertions on your own logic reliable.