Deepagents tracks real session cost but only shows a toast. Nothing stops the loop from spending past a limit. A before_model hook that can jump to "end" closes that gap.
Past ~20-30 tools, sending every schema every turn hurts cost and accuracy. A lexical scorer plus a registry-search hatch fixes most of it — no embeddings needed.
Most agent frameworks observe model calls and allow rewriting them only after they reach the model, making an understanding of callbacks and middleware essential.