The Warning That Never Stops the Agent
Deepagents tracks real session cost but only shows a toast. Nothing stops the loop from spending past a limit. A before_model hook that can jump to "end" closes that gap.
Join the DZone community and get the full member experience.
Join For FreeI went looking for a flag that could cap the amount spent on a session on DeepAgents. The kind of hard spending cap that stops an agent loop before it burns through a budget. What exists instead is a warning, and the gap between "warning" and "stopping" turned out to be the more interesting story.
What Already Exists: A Number With No Teeth
deepagents-code has a CostTrackingMiddleware that owns a thread's cumulative spend. This is a real, checkpointed dollar figure priced from actual per-request token usage, not an estimate. A separate feature built on top of that number, merged earlier, adds a one-time warning once the total crosses a configured threshold (default $50):
# app.py, roughly what ships today
threshold = self._session_cost_warning_threshold_usd
if (
not self._session_cost_warning_shown
and 0 < threshold < self._session_cost_usd
):
self._session_cost_warning_shown = True
self.notify(
f"Estimated session cost is {format_cost(self._session_cost_usd)}, "
f"above the configured {format_cost(threshold)} threshold. Consider "
"/offload to reduce context usage or /clear to start fresh.",
title="Session cost warning",
severity="warning",
timeout=12,
markup=False,
)
Read that once more: It's a self.notify(...) call. A toast! The agent's own loop has no idea this happened. Nothing in the code above touches the graph, the model call, or the next tool invocation. If you're watching the terminal, you see the warning and can intervene by hand (/offload, /clear, or just killing the process). If you're not watching and it's a headless CI run — for example, an unattended overnight session or a tool-call loop that's quietly retrying against a flaky API — this number keeps climbing, and nothing stops it.
That's the gap: a cost number that can only ever inform a human, never the loop that's actually spending the money.
Why the Fix Isn't "Just Check the Number Somewhere"
The obvious instinct is to add a if cost > limit: stop check. The real work is in where that check has to live and what "stop" has to mean to a LangGraph agent loop.
Where: The check needs to run before the next model call, not after. Checking after a call has already happened is too late to prevent its cost. LangGraph gives middleware a before_model hook for exactly this. LangChain's own ModelCallLimitMiddleware (a call-count limiter, not a cost one) already establishes the pattern: check a condition in before_model, and if it's tripped, return an update that redirects the graph to end instead of letting the model call happen.
What "stop" means: A plain return None from before_model just lets the loop continue. To actually halt, the hook needs @hook_config(can_jump_to=["end"]) and has to return {"jump_to": "end", ...} which is a real graph-control signal, not a value the caller has to notice and act on:
# cost_tracking.py — the actual hook, as committed
@hook_config(can_jump_to=["end"])
def before_model(
self,
state: CostState,
runtime: Runtime[ContextT],
) -> dict[str, Any] | None:
"""Halt the run before the next model call if the hard cost cap is met.
Checked against the checkpointed cumulative total from the *previous*
step -- the same figure the TUI's soft warning reads via
`_set_session_cost` -- so this fires at the same point in the loop a
user would already have seen the warning toast, just before the next
request that would push spend further over the configured cap.
"""
if self._nested or self._hard_limit_usd is None:
return None
total_usd = state.get("_session_cost_usd")
if (
isinstance(total_usd, bool)
or not isinstance(total_usd, int | float)
or not math.isfinite(total_usd)
):
return None
if total_usd < self._hard_limit_usd:
return None
if self._exit_behavior == "error":
raise CostLimitExceededError(total_usd, self._hard_limit_usd)
limit_message = _build_cost_limit_message(total_usd, self._hard_limit_usd)
return {"jump_to": "end", "messages": [AIMessage(content=limit_message)]}
Two details worth calling out, because they're the kind of thing that looks like overcaution until you hit the failure it's guarding against:
isinstance(total_usd, bool) before the numeric check. In Python, bool is a subclass of int, so isinstance(True, int | float) is True and True < 5.0 evaluates fine (True == 1). Without the explicit bool guard, a stray True sitting in a state field meant for a float would silently be treated as $1.00 and could either falsely trip the halt or falsely pass it, depending on the cap. Cheap to guard against, expensive to debug if you don't.
Checked only on the non-nested instance. CostTrackingMiddleware also runs on subagents, where its _session_cost_usd channel tracks that subagent's own local spend before it's transferred back to the parent's running total. Checking the hard cap there would be checking the wrong number - a subagent's small local total against a cap meant to bound the whole session's spend. self._nested gates this out entirely.
The Part That Isn't the Algorithm: Getting the Number Across a Process Boundary
Here's what made this bigger than a one-file change: dcode doesn't run the agent loop in the same process as the CLI. It starts a LangGraph server in a subprocess and talks to it over langgraph-sdk. Every configuration value the agent needs, like model name, sandbox type, recursion limit, and now this cap, has to survive that boundary, which in this codebase means round-tripping through environment variables on a ServerConfig dataclass:
# _server_config.py
max_cost_usd: float | None = None
"""Explicit hard cap, in USD, on the main thread's cumulative estimated cost."""
def to_env(self) -> dict[str, str | None]:
return {
...
"MAX_COST_USD": (
str(self.max_cost_usd) if self.max_cost_usd is not None else None
),
...
}
@classmethod
def from_env(cls) -> ServerConfig:
return cls(
...
max_cost_usd=_read_env_float("MAX_COST_USD", default=None),
...
)
_read_env_float didn't exist before this - every other numeric config value in this file is an int (recursion_limit, turn counts), so there was no float-reading helper to reuse. One new function, matching the existing _read_env_int's shape exactly, and the round-trip works the same way every other config value already does.
The full path, in order: a --max-cost CLI flag → a resolver that checks the flag, then a config.toml entry, then "disabled" → create_cli_agent(max_cost_usd=...) → ServerConfig → serialized to an environment variable → the subprocess reads it back → create_cli_agent again, this time inside the subprocess → CostTrackingMiddleware(hard_limit_usd=...). Six hops for one float, and every one of them was necessary - skip the ServerConfig round-trip and the flag works in a unit test but silently does nothing the moment you actually run dcode, because the subprocess that runs the real agent loop never sees it.

Verifying It Live
Unit tests covering the before_model logic in isolation are necessary but not sufficient here. They'd pass even if one of those six hops silently dropped the value, because a unit test calls the middleware directly and never exercises the subprocess boundary at all. The only way to know the flag actually works is to run the real CLI against a real model:
$ dcode -n "Read sample.txt, then read it again, then read it a third time. \
Do this as three separate tool calls, one per turn." \
--model anthropic:claude-haiku-4-5 \
--max-cost 0.0001 \
--max-turns 6
I'll read sample.txt, then read it again on the next turn, then a third
time on the turn after that.
First read:
Calling tool: read_file
Session halted: estimated cost $0.02 has reached the configured limit
of <$0.01. Raise the limit (e.g. `--max-cost`) or start a new session
to continue.
Task completed
Usage Stats
Provider Model Reqs InputTok OutputTok Cost
anthropic claude-haiku-4-5 1 13.6K 147 $0.02
The model was told to make three tool calls, one per turn. It made exactly one. The middleware checkpointed that turn's real cost (0.02), the next beforemodel check found the cumulative total over the(deliberately absurd) 0.0001 cap, and the graph jumped to end with the injected message instead of continuing to spend on turns two and three. Total cost of proving this worked: two cents.
Why This Is Worth a Hard Stop and Not Just a Bigger Warning
You could imagine closing this gap by making the warning louder: repeat it every turn instead of once, or block user input until it's acknowledged. That doesn't fix the actual failure mode, which is specifically the unattended case. A louder toast is still a toast. Nothing short of a return value the graph itself has to obey closes that gap, which is why the fix has to live in before_model, not in the terminal UI layer where the existing warning already sits.
It's also worth being honest about what this doesn't fix: the check runs before a model call, using the cost checkpointed from the previous one. A single turn can still overshoot the cap if the cap is 5.00 and the agent is at 4.99. The next call still happens in full and might land at 6.00 before the halt fires on the turn after. That's not a bug so much as an inherent property of checking after the fact rather than metering mid-request, and it's the same tradeoff that ModelCallLimitMiddleware makes for call counts. A cap is a backstop against runaway, unattended spend. It's not a precise billing guarantee down to the last cent.
Takeaways
The general lesson here isn't really about cost. It's that a warning and a limit are two different features wearing the same clothing, and it's easy to ship the first while believing you've shipped the second. The warning reads the same number, uses the same word ("threshold"), and looks like it's doing the same job right up until someone isn't in the room to read it. The tell is always the same: does the check return a value the system has to act on, or does it just call something with "notify" in the name?
Second, a fix that only works in-process is only half a fix once your architecture has a subprocess boundary in it. The six-hop threading here wasn't extra caution, but it was the actual scope of the problem, and skipping any one hop would have shipped a flag that silently does nothing.
Third, if you can run the real thing end-to-end for two cents, there's no good reason to trust a mock's word for whether a fix actually works.
Opinions expressed by DZone contributors are their own.
Comments