Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
Temporal’s durable orchestration and LangGraph checkpoints enable research agents to recover from worker crashes and resume from saved progress.
Join the DZone community and get the full member experience.
Join For FreeA research agent can spend minutes or hours accumulating search results, model outputs, citations, and intermediate synthesis. Losing that work because a Python worker restarts is not an acceptable failure mode. The useful design target is therefore not an immortal worker process but a disposable worker whose completed research remains durable.
Temporal’s official LangGraph integration provides that boundary: LangGraph still defines agent state and control flow, while Temporal supplies durable execution, Activity retries, timeouts, and replay-oriented recovery. As of September 2026, the integration is in Public Preview; its Python API is still marked experimental, and it requires temporalio 1.27.0 or later.
Make the Process Disposable, Not the State
Temporal treats a Workflow Execution as durable state rather than as the lifetime of one worker process. After a crash or infrastructure interruption, execution can be reconstructed and resumed; Temporal explicitly describes its execution model as surviving crashes, network failures, and infrastructure outages. That property changes how LangGraph persistence should be approached. In a conventional LangGraph deployment, a checkpointer persists thread-scoped graph state, and InMemorySaver explicitly loses checkpoints when the process restarts. Under the Temporal integration, however, Temporal owns execution durability; the integration documentation specifically recommends InMemorySaver only when LangGraph itself requires a checkpointer, such as for interrupts, rather than adding PostgreSQL or Redis merely to duplicate Workflow durability.
That separation matters because two persistence systems for the same execution create two authorities for progress. A database checkpoint may represent one view of graph advancement while Temporal history represents another. The cleaner model makes Temporal’s Workflow history the execution ledger and keeps LangGraph state suitable for durable orchestration. External storage still has a role for application data, documents, embeddings, and cross-run memory, but not as a second scheduler checkpoint for the same run. Temporal’s integration makes this ownership explicit by assigning execution durability to Temporal while retaining LangGraph’s graph semantics.
Put Uncertainty Behind Activities
The integration requires every LangGraph node or Functional API task to declare where it executes. Network calls, LLM requests, database access, file I/O, current-time reads, randomness, and other nondeterministic operations belong in Temporal Activities. Pure state transformations and lightweight routing can run inside the Workflow, where replay requires deterministic behavior. Activity nodes receive Temporal timeout and retry policies; the plugin rejects LangGraph retry policies for these nodes and expects Temporal RetryPolicy configuration instead.
A research graph can make that boundary explicit without changing the overall LangGraph programming model:
research = StateGraph(ResearchState)
research.add_node(
"search",
search_sources,
metadata={
"execute_in": "activity",
"start_to_close_timeout": timedelta(minutes=2),
"retry_policy": RetryPolicy(maximum_attempts=4),
},
)
research.add_node(
"merge",
merge_evidence,
metadata={"execute_in": "workflow"},
)
research.add_edge(START, "search")
research.add_edge("search", "merge")
The search node is an Activity because its result depends on an external service and may fail transiently. The merge node can remain in the Workflow when it only combines already-recorded evidence. This is the central recoverability rule: expensive uncertainty crosses an Activity boundary; deterministic orchestration stays replayable. Conditional-edge functions deserve the same discipline because the integration always runs them in the Workflow and requires them to be deterministic and asynchronous.
Let Temporal Become the Execution Ledger
The graph becomes available to Workflows through LangGraphPlugin, while execution occurs on a normal Temporal Worker. Shared timeout and retry defaults can live on the plugin, but execute_in still has to be declared on each node so that an accidental nondeterministic Workflow node cannot hide behind a global default.
plugin = LangGraphPlugin(
graphs={"research": research},
default_activity_options={
"start_to_close_timeout": timedelta(minutes=2),
"retry_policy": RetryPolicy(maximum_attempts=4),
},
)
worker = Worker(
client,
task_queue="research-agent",
workflows=[ResearchWorkflow],
plugins=[plugin],
)
await worker.run()
Inside the Workflow, invoking the registered graph remains small:
app = graph("research").compile()
result = await app.ainvoke({
"query": request.query,
"evidence": [],
"draft": None,
})
If the worker disappears after an Activity has completed and its outcome has become part of durable execution, recovery does not depend on reconstructing arbitrary Python process memory. If an Activity attempt is interrupted by a worker crash or transient failure, Temporal can retry that Activity according to its policy rather than restarting the entire research graph. Temporal’s LangGraph documentation explicitly identifies worker crashes as a condition under which an Activity-wrapped node can rerun, while Temporal’s platform model preserves Workflow execution across process and infrastructure failures.
Retries also change the design of external writes. Any Activity whose side effect can survive a failed attempt should carry a stable operation key so a retry resolves to the same logical write rather than creating another one. This follows directly from the fact that Activity attempts can rerun after failures. A persistence Activity can use the run and revision as that identity:
async def save_draft(state: ResearchState) -> dict:
operation_id = f"{state['run_id']}:{state['revision']}"
await repository.upsert(operation_id, state["draft"])
return {"saved_revision": state["revision"]}
The retry-safe behavior in this example comes from the storage contract around operation_id, not from assuming that an Activity executes only once.
Keep Memory and Progress Separate
LangGraph’s normal persistence model distinguishes thread checkpoints from longer-lived stores. Temporal narrows that choice inside this integration. The documentation states that LangGraph Store objects are not available inside Activity-wrapped nodes because live store state cannot cross the Activity boundary and an Activity may execute on a different worker. Per-run memory therefore belongs in Workflow or graph state. Shared memory across runs belongs in an external database that every relevant worker can access independently.
This distinction also prevents a subtle mistake around InMemorySaver. LangGraph documentation correctly warns that InMemorySaver is process-local and disappears on restart. Temporal’s integration can still recommend it for LangGraph features such as interrupt() because it is not being asked to provide crash durability there; Temporal is. The two statements are compatible once the ownership boundary is explicit.
Streaming needs similar treatment. Activity-side LangGraph streaming has at-least-once delivery semantics per Activity attempt. A failed attempt can publish chunks that appear again when the Activity retries because earlier stream writes are not rolled back. Progress streams should therefore be treated as advisory presentation data or deduplicated with sequence identifiers. Temporal’s documentation explicitly recommends deduplication or relying on the final Workflow result as authoritative state.
Keep Long Research From Becoming Long History
A deep-research agent may loop through dozens of retrieval, critique, and synthesis passes. Temporal records execution events, so sufficiently long graphs can eventually approach Workflow history limits. The LangGraph integration addresses this with Continue-As-New plus a task-result cache. cache() returns a serializable representation of completed LangGraph task results; passing that cache into graph(..., cache=...) in the new run allows previously completed nodes to avoid re-execution.
app = graph("research", cache=prior_cache).compile()
state = await app.ainvoke(state)
workflow.continue_as_new(
args=[state, cache()]
)
The transition should occur at a deliberate research boundary before history becomes excessive. Continue-As-New starts fresh execution history while application state and the LangGraph result cache can be carried into the successor run. The distinction is important: a new Temporal run does not need to imply a new research effort. Previously completed node results remain reusable even as the accumulated history is reset for continued execution.
Recovery Is the Feature
A recoverable research agent is produced by assigning durability to the correct layer, not by trying to keep a worker alive forever. LangGraph should express the agent’s graph and state transitions; Temporal should own replay-oriented recovery, retries, timeout enforcement, and durable execution; external databases should hold genuinely shared application memory and retry-safe side effects.
With nondeterministic work isolated in Activities, deterministic control flow kept inside the Workflow, streams treated as non-authoritative, and Continue-As-New used for long histories, a worker can be terminated at an inconvenient moment without turning completed research into lost work. The process becomes replaceable, while the research remains part of the durable execution.
Opinions expressed by DZone contributors are their own.
Comments