Building Production-Ready AI Agents with LangGraph
The gap between a LangGraph demo and a production agent is almost entirely about state, persistence, and interrupts — not the model doing the reasoning.
By Naeem Akhtar · 9 min read
Why a graph instead of a chain
A linear prompt chain works fine until an agent needs to branch, retry a failed step, loop back for clarification, or pause and wait for a person. At that point, forcing the logic through a chain means bolting conditional hacks onto something that was never designed to hold state across a non-linear path.
LangGraph's core idea is to model the agent explicitly as a directed graph: nodes are units of work, edges define what can follow what, and a typed state object is threaded through and updated at every node. That structure is what makes branching, retries, and human review checkpoints first-class citizens instead of workarounds.
The checkpointer decision comes before anything else
Every LangGraph deployment persists its state through a checkpointer, and the choice of backend is the single most consequential production decision in the whole build. In-memory checkpointing is fine for local development, but it means agent state — including anything mid-approval or mid-conversation — disappears the moment the process restarts.
from langgraph.checkpoint.postgres import PostgresSaver
with PostgresSaver.from_conn_string(DATABASE_URL) as checkpointer:
checkpointer.setup()
graph = builder.compile(checkpointer=checkpointer)For production, that means a persistent backend — PostgresSaver is the common choice for teams already running Postgres — with connection pooling configured for the concurrency you actually expect, and a deliberate retention policy for checkpoint data rather than storing it indefinitely by default.
Designing interrupts for human-in-the-loop
The pattern that makes async approval workflows possible is the interrupt: a node can pause graph execution mid-run, surface a payload for a human to review, and resume later from that exact point once a decision comes back — without replaying the conversation or losing context.
def send_external_email(state):
decision = interrupt({
"action": "send_email",
"to": state["recipient"],
"draft": state["draft_body"],
})
if decision["approved"]:
return {"status": "sent"}
return {"status": "rejected", "reason": decision.get("reason")}This matters because a real approval can take minutes or days — a reviewer might be offline, or the action might need sign-off from someone in a different timezone. Treating the interrupt as a first-class paused state, not an exception to catch and retry, is what makes that wait safe.
Streaming and tracing aren't optional at production scale
Two things become necessary the moment an agent handles real traffic instead of a demo script: streaming intermediate output to whatever's consuming it (a UI, a voice pipeline, a chat interface), so users aren't staring at a spinner for a multi-step task; and tracing every node transition and tool call, so when an agent does something unexpected, the team can see exactly why — not just that it happened.
- Log every node entry/exit, the state diff at each step, and every tool call with its arguments and result.
- Stream partial output where the latency of a full multi-step run would otherwise feel broken to the end user.
- Capture interrupt events distinctly from errors — a paused-for-review state is expected behavior, not a failure, and conflating the two in logs makes debugging much harder.
The four pillars of a production deployment
| Pillar | What it covers |
|---|---|
| Persistent checkpointing | State survives restarts and long-running, multi-day approval waits. |
| Tracing and observability | Every node transition and tool call is visible, not just the final output. |
| Interrupt-based governance | Human review checkpoints are a first-class graph state, not an exception handler. |
| A real deployment target | A managed platform or a self-hosted, containerized service built to run the graph continuously, not a notebook. |
Skipping any one of these tends to work fine right up until it doesn't — usually the first time a process restarts mid-approval, or the first time someone asks why the agent did what it did.
Mistakes that only show up under real load
- No idempotency on tool calls — a retry after a timeout re-executes a side-effecting action (like sending an email) a second time.
- Never testing checkpoint recovery — the first real test of "does this resume correctly" happens during an actual production incident instead of in staging.
- Treating interrupts as exceptional rather than expected, which makes monitoring noisy and hides real errors in a sea of paused-for-review alerts.
- No retention policy on checkpoint data, which quietly accumulates PII and becomes a compliance problem long before anyone notices the storage cost.
Frequently asked questions
Do I need PostgresSaver for a LangGraph agent, or is in-memory checkpointing enough?
In-memory checkpointing is fine for local development and demos, but any production deployment needs a persistent checkpointer such as PostgresSaver. Without it, agent state — including anything paused mid-approval — is lost the moment the process restarts, which is not acceptable once real users or business processes depend on the agent.
How does human-in-the-loop actually work in LangGraph?
A node calls an interrupt with a payload describing the pending decision, which pauses graph execution at that exact point. The graph resumes later — potentially minutes or days afterward — from that same state once a human decision comes back, without needing to replay or reconstruct the prior conversation.
What's the difference between LangGraph and a standard prompt chain?
A prompt chain is linear — one step feeds the next. LangGraph models the agent as a graph of nodes and edges with typed state threaded through it, which makes branching, retries, loops, and human review checkpoints native parts of the design instead of workarounds bolted onto a linear structure.
Is LangSmith required for LangGraph production deployments?
Not strictly required, but some form of tracing that logs every node transition and tool call is effectively mandatory once an agent affects real customers or money. LangSmith is the common managed option; a self-hosted logging and tracing setup can serve the same purpose if that's the team's preference.