# Deploying, Observability and Guardrails — LangGraph Agents & Multi-Agent Systems

Source: https://www.geekswithgeeks.com/en/langgraph-agents/t-deploy

> Run graphs safely in production.

## Durable, observable, bounded

In production, run graphs with a **durable checkpointer** (Postgres or similar), behind an API that authenticates users and derives `thread_id` from the session; use **background workers or a managed agent server** for long runs, with timeouts and cancellation. Add **tracing** for every run (inputs, node outputs, tool calls, tokens, latency, errors) with sensitive data redacted, plus metrics and alerts on failure rate, loop-cap hits, cost per run and approval backlog. **Guardrails**: least-privilege tools, argument validation, allow-listed actions, human approval for risky steps, output filters, treating retrieved text and tool output as untrusted data (prompt injection), rate limits, and an easy kill switch. Version the graph and prompts, and roll out changes gradually.

## A production checklist

Use it as a review list before launch.

```text
[ ] durable checkpointer; thread_id derived from authenticated session
[ ] recursion limit + attempt counters + time and cost budget per run
[ ] tools: least privilege, validated args, allow-listed actions, idempotent writes
[ ] human approval for irreversible / costly actions (interrupt) + audit record
[ ] retrieved text and tool output treated as untrusted data
[ ] tracing with redaction; alerts on failures, loop caps, cost, stuck approvals
[ ] evaluation suite (outcome + trajectory + safety) runs in CI
[ ] versioned graph/prompts, staged rollout, kill switch
```

**Quiz:** Where should the thread_id come from in production?

- [ ] The model name
- [ ] Chosen by the client freely
- [ ] Always the same constant
- [x] Derived from the authenticated user and session

*Answer:* Derived from the authenticated user and session. Otherwise one user could load another user's conversation.
