Lesson 28 / 31

Versions, Cost, Latency and Security

Operate a framework-based app safely and affordably.

Pin, measure, protect

Versions: both libraries evolve quickly and have been reorganised into many packages, so pin exact versions in a lockfile, read release notes before upgrading, and keep tests that catch breakages. Cost and latency: log token usage per call, cap output tokens, retrieve fewer and better chunks, cache repeated work (embeddings, identical queries, prompt prefixes), stream responses, and use a smaller model for easy steps. Security: keep API keys in environment variables or a secrets manager, never in code or prompts; apply access control inside retrieval; treat retrieved text as untrusted; redact sensitive data before logging or tracing; review third-party loaders and integrations, since they run code with your permissions.

A pinned requirements file (illustrative)

Pin the versions you tested with. The numbers shown are the versions used to run this course's examples; choose versions that suit your project.

# requirements.txt
langchain-core==1.6.6
langchain-text-splitters==1.1.2
llama-index-core==0.14.25
# add provider packages (e.g. langchain-openai) with exact versions too

Quick check: Why pin library versions?

  • Pinning makes the model smarter
  • These libraries change quickly and upgrades can break code
  • It is required by Python
  • It reduces token cost
Answer

These libraries change quickly and upgrades can break code — Stable versions make behaviour reproducible until you choose to upgrade.