Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Code Review Core Practices
Getting Started With DevSecOps
I went looking for a flag that could cap the amount spent on a session on DeepAgents. The kind of hard spending cap that stops an agent loop before it burns through a budget. What exists instead is a warning, and the gap between "warning" and "stopping" turned out to be the more interesting story. What Already Exists: A Number With No Teeth deepagents-code has a CostTrackingMiddleware that owns a thread's cumulative spend. This is a real, checkpointed dollar figure priced from actual per-request token usage, not an estimate. A separate feature built on top of that number, merged earlier, adds a one-time warning once the total crosses a configured threshold (default $50): Python # app.py, roughly what ships today threshold = self._session_cost_warning_threshold_usd if ( not self._session_cost_warning_shown and 0 < threshold < self._session_cost_usd ): self._session_cost_warning_shown = True self.notify( f"Estimated session cost is {format_cost(self._session_cost_usd)}, " f"above the configured {format_cost(threshold)} threshold. Consider " "/offload to reduce context usage or /clear to start fresh.", title="Session cost warning", severity="warning", timeout=12, markup=False, ) Read that once more: It's a self.notify(...) call. A toast! The agent's own loop has no idea this happened. Nothing in the code above touches the graph, the model call, or the next tool invocation. If you're watching the terminal, you see the warning and can intervene by hand (/offload, /clear, or just killing the process). If you're not watching and it's a headless CI run — for example, an unattended overnight session or a tool-call loop that's quietly retrying against a flaky API — this number keeps climbing, and nothing stops it. That's the gap: a cost number that can only ever inform a human, never the loop that's actually spending the money. Why the Fix Isn't "Just Check the Number Somewhere" The obvious instinct is to add a if cost > limit: stop check. The real work is in where that check has to live and what "stop" has to mean to a LangGraph agent loop. Where: The check needs to run before the next model call, not after. Checking after a call has already happened is too late to prevent its cost. LangGraph gives middleware a before_model hook for exactly this. LangChain's own ModelCallLimitMiddleware (a call-count limiter, not a cost one) already establishes the pattern: check a condition in before_model, and if it's tripped, return an update that redirects the graph to end instead of letting the model call happen. What "stop" means: A plain return None from before_model just lets the loop continue. To actually halt, the hook needs @hook_config(can_jump_to=["end"]) and has to return {"jump_to": "end", ...} which is a real graph-control signal, not a value the caller has to notice and act on: Python # cost_tracking.py — the actual hook, as committed @hook_config(can_jump_to=["end"]) def before_model( self, state: CostState, runtime: Runtime[ContextT], ) -> dict[str, Any] | None: """Halt the run before the next model call if the hard cost cap is met. Checked against the checkpointed cumulative total from the *previous* step -- the same figure the TUI's soft warning reads via `_set_session_cost` -- so this fires at the same point in the loop a user would already have seen the warning toast, just before the next request that would push spend further over the configured cap. """ if self._nested or self._hard_limit_usd is None: return None total_usd = state.get("_session_cost_usd") if ( isinstance(total_usd, bool) or not isinstance(total_usd, int | float) or not math.isfinite(total_usd) ): return None if total_usd < self._hard_limit_usd: return None if self._exit_behavior == "error": raise CostLimitExceededError(total_usd, self._hard_limit_usd) limit_message = _build_cost_limit_message(total_usd, self._hard_limit_usd) return {"jump_to": "end", "messages": [AIMessage(content=limit_message)]} Two details worth calling out, because they're the kind of thing that looks like overcaution until you hit the failure it's guarding against: isinstance(total_usd, bool) before the numeric check. In Python, bool is a subclass of int, so isinstance(True, int | float) is True and True < 5.0 evaluates fine (True == 1). Without the explicit bool guard, a stray True sitting in a state field meant for a float would silently be treated as $1.00 and could either falsely trip the halt or falsely pass it, depending on the cap. Cheap to guard against, expensive to debug if you don't. Checked only on the non-nested instance. CostTrackingMiddleware also runs on subagents, where its _session_cost_usd channel tracks that subagent's own local spend before it's transferred back to the parent's running total. Checking the hard cap there would be checking the wrong number - a subagent's small local total against a cap meant to bound the whole session's spend. self._nested gates this out entirely. The Part That Isn't the Algorithm: Getting the Number Across a Process Boundary Here's what made this bigger than a one-file change: dcode doesn't run the agent loop in the same process as the CLI. It starts a LangGraph server in a subprocess and talks to it over langgraph-sdk. Every configuration value the agent needs, like model name, sandbox type, recursion limit, and now this cap, has to survive that boundary, which in this codebase means round-tripping through environment variables on a ServerConfig dataclass: Python # _server_config.py max_cost_usd: float | None = None """Explicit hard cap, in USD, on the main thread's cumulative estimated cost.""" def to_env(self) -> dict[str, str | None]: return { ... "MAX_COST_USD": ( str(self.max_cost_usd) if self.max_cost_usd is not None else None ), ... } @classmethod def from_env(cls) -> ServerConfig: return cls( ... max_cost_usd=_read_env_float("MAX_COST_USD", default=None), ... ) _read_env_float didn't exist before this - every other numeric config value in this file is an int (recursion_limit, turn counts), so there was no float-reading helper to reuse. One new function, matching the existing _read_env_int's shape exactly, and the round-trip works the same way every other config value already does. The full path, in order: a --max-cost CLI flag → a resolver that checks the flag, then a config.toml entry, then "disabled" → create_cli_agent(max_cost_usd=...) → ServerConfig → serialized to an environment variable → the subprocess reads it back → create_cli_agent again, this time inside the subprocess → CostTrackingMiddleware(hard_limit_usd=...). Six hops for one float, and every one of them was necessary - skip the ServerConfig round-trip and the flag works in a unit test but silently does nothing the moment you actually run dcode, because the subprocess that runs the real agent loop never sees it. Verifying It Live Unit tests covering the before_model logic in isolation are necessary but not sufficient here. They'd pass even if one of those six hops silently dropped the value, because a unit test calls the middleware directly and never exercises the subprocess boundary at all. The only way to know the flag actually works is to run the real CLI against a real model: Shell $ dcode -n "Read sample.txt, then read it again, then read it a third time. \ Do this as three separate tool calls, one per turn." \ --model anthropic:claude-haiku-4-5 \ --max-cost 0.0001 \ --max-turns 6 I'll read sample.txt, then read it again on the next turn, then a third time on the turn after that. First read: Calling tool: read_file Session halted: estimated cost $0.02 has reached the configured limit of <$0.01. Raise the limit (e.g. `--max-cost`) or start a new session to continue. Task completed Usage Stats Provider Model Reqs InputTok OutputTok Cost anthropic claude-haiku-4-5 1 13.6K 147 $0.02 The model was told to make three tool calls, one per turn. It made exactly one. The middleware checkpointed that turn's real cost (0.02), the next beforemodel check found the cumulative total over the(deliberately absurd) 0.0001 cap, and the graph jumped to end with the injected message instead of continuing to spend on turns two and three. Total cost of proving this worked: two cents. Why This Is Worth a Hard Stop and Not Just a Bigger Warning You could imagine closing this gap by making the warning louder: repeat it every turn instead of once, or block user input until it's acknowledged. That doesn't fix the actual failure mode, which is specifically the unattended case. A louder toast is still a toast. Nothing short of a return value the graph itself has to obey closes that gap, which is why the fix has to live in before_model, not in the terminal UI layer where the existing warning already sits. It's also worth being honest about what this doesn't fix: the check runs before a model call, using the cost checkpointed from the previous one. A single turn can still overshoot the cap if the cap is 5.00 and the agent is at 4.99. The next call still happens in full and might land at 6.00 before the halt fires on the turn after. That's not a bug so much as an inherent property of checking after the fact rather than metering mid-request, and it's the same tradeoff that ModelCallLimitMiddleware makes for call counts. A cap is a backstop against runaway, unattended spend. It's not a precise billing guarantee down to the last cent. Takeaways The general lesson here isn't really about cost. It's that a warning and a limit are two different features wearing the same clothing, and it's easy to ship the first while believing you've shipped the second. The warning reads the same number, uses the same word ("threshold"), and looks like it's doing the same job right up until someone isn't in the room to read it. The tell is always the same: does the check return a value the system has to act on, or does it just call something with "notify" in the name? Second, a fix that only works in-process is only half a fix once your architecture has a subprocess boundary in it. The six-hop threading here wasn't extra caution, but it was the actual scope of the problem, and skipping any one hop would have shipped a flag that silently does nothing. Third, if you can run the real thing end-to-end for two cents, there's no good reason to trust a mock's word for whether a fix actually works.
Multi-step flows are common in UX, including onboarding, checkout, account setup, approval processes, configuration wizards, and administrative tasks. These require users to move through multiple screens while continuing a consistent working state. The challenge is to keep this state active for the duration of the interaction, but not beyond. Request scope is too short, while session scope often extends longer than the business process needs. Jakarta Faces handles this with @FlowScoped, which manages state based on the lifecycle of a flow instead of a single page or the entire session. This article uses a customer segmentation application to demonstrate how a flow can guide users through configuration, preview, and confirmation, while maintaining state across each step. This approach creates a cleaner model for wizard-style UX: the scope begins when the user enters the flow, persists during navigation, and ends upon exit. Why Jakarta Faces Still Matters Jakarta Faces continues to be relevant because many enterprise applications are developed and maintained by teams with strong Java expertise. In these environments, a server-side UI framework limits context switching, keeps validation and navigation close to the application model, and allows teams to reuse the same language, dependency injection model, and enterprise APIs throughout the stack. Architecturally, if the team is proficient in Java and the application is form-driven, workflow-oriented, or back-office focused, introducing a separate frontend stack does not necessarily offer an advantage. Component libraries such as PrimeFaces further support this approach. Rather than building tables, dialogs, forms, charts, wizards, and validation from scratch, teams can use reusable components within the Jakarta EE programming model. This can accelerate delivery and lessen the need for custom frontend infrastructure. While the decision should be based on product and team context, for Java-focused enterprise teams, Jakarta Faces is a pragmatic architectural choice, not just a legacy option. A few publicly documented examples of organizations that have used Jakarta EE/JSF or PrimeFaces include: NASAWalmart LabsLufthansaRakutenCommerzbankUnited NationsPenn State UniversityHarvard UniversityTelefonicaBig LotsComfortel (telecommunications)Various commercial banks and financial institutions Overview of Jakarta Faces Scopes Jakarta Faces applications maintain managed bean state for varying durations. Choosing the right scope is an architectural decision and must match the user interaction's lifetime. Some state is limited to a single HTTP request, a single page, a multi-step flow, or the entire user session. @RequestScoped is the shortest-lived scope. It creates a bean instance for each HTTP request and discards it when the request completes. It suits stateless actions, simple submissions, and operations that don't need to continue across navigation or Ajax interactions.@ViewScoped retains the bean while the user stays on the same Faces view. It is ideal for pages with forms, tables, filtering, pagination, dialogs, or Ajax interactions that update the same page multiple times. The state is discarded when the user navigates to a different view.@FlowScoped is intended for business interactions spanning multiple views, such as checkout, onboarding, configuration wizards, approval workflows, or customer segmentation. The bean remains active throughout the flow and is destroyed when the flow ends. Its lifecycle falls between view scope and session scope.@SessionScoped maintains state for the entire user session. It suits information that remains across multiple pages, such as user preferences or session-level context. Avoid using it for temporary workflow state, as this can unnecessarily extend the state’s lifetime.@ApplicationScoped has the broadest lifetime, sharing a single bean instance across the entire application and all users. It suits shared services, caches, configuration, or application-wide state, but not per-user or per-flow data unless explicitly designed for sharing and thread safety. Building a Multi-Step Experience With @FlowScoped This article focuses on Jakarta Faces flow, which models user engagements spanning multiple pages but shorter than a full HTTP session. In the customer-segmentation example, the administrator configures thresholds, previews their impact, reviews the configuration, and starts the operation. These steps form a single business process and should share the same state. Jakarta Faces represents this process with a flow definition and a flow-scoped managed bean. In this project, the flow resides in the customer-segmentation directory: reStructuredText src/main/webapp/ └── customer-segmentation/ ├── customer-segmentation-flow.xml ├── configure.xhtml ├── preview.xhtml └── review.xhtml Aligning the directory, flow identifier, and bean name clarifies their relationship. Here, the flow stays named customer-segmentation, defined in customer-segmentation/customer-segmentation-flow.xml, and the Java bean uses @FlowScoped("customer-segmentation"). The value passed to @FlowScoped ties the bean’s lifecycle to the corresponding Faces flow. The XML file defines the flow’s structure and navigation boundaries, but does not store business state. In this example, configure is the starting point, followed by preview and review. Each view has an identifier and references its corresponding XHTML document: XML <flow-definition id="customer-segmentation"> <start-node>configure</start-node> <view id="configure"> <vdl-document> /customer-segmentation/configure.xhtml </vdl-document> </view> <view id="preview"> <vdl-document> /customer-segmentation/preview.xhtml </vdl-document> </view> <view id="review"> <vdl-document> /customer-segmentation/review.xhtml </vdl-document> </view> <flow-return id="home"> <from-outcome>/index?faces-redirect=true</from-outcome> </flow-return> </flow-definition> The start-node specifies where the interaction begins. The <view> elements define the flow’s pages, and <flow-return> determines how the application exits. Returning the outcome home ends the flow, redirects the user to the dashboard, and discards the flow-scoped state. Navigation within the flow preserves the state, while exiting ends the conversation. On the Java side, CustomerSegmentationFlow manages the conversation state as both a named Faces bean and a flow-scoped bean: Java @Named @FlowScoped("customer-segmentation") public class CustomerSegmentationFlow implements Serializable { @Inject private CustomerSegmentationFlowService flowService; private CustomerSegmentationFlowState state; @PostConstruct public void initialize() { state = flowService.initializeState(); } // ... } @Named makes the bean accessible to Faces pages through Expression Language, while @FlowScoped("customer-segmentation") assigns its lifecycle to the specific flow. The state created during @PostConstruct remains across requests and page transitions during the flow, so CustomerSegmentationFlowState remains available throughout the interaction. This distinction sets @FlowScoped apart from @ViewScoped. With view scope, moving from configure.xhtml to preview.xhtml starts a new conversation. With flow scope, both views remain part of the same business interaction. The scope follows the conversation, not individual pages. Navigation methods on the bean then return outcomes that correspond to nodes defined by the flow: Java public String preview() { flowService.preview(state); return "preview"; } public String review() { return "review"; } Returning "preview" moves the user to the preview view in the flow definition, while "review" advances to the next view. Since both destinations are within customer-segmentation, the same flow-scoped bean and its state remain active. The final action demonstrates the other side of the lifecycle: Java public String execute() { long executionId = flowService.start(state); // message handling omitted return "home"; } Home is not a page within the wizard. It corresponds to the <flow-return id="home"> element in the XML definition. When this outcome occurs, Faces exits the flow, redirects to the dashboard, and discards the associated state. @FlowScoped is ideal for processes such as checkout, onboarding, registration, approval, configuration, and administrative wizards. It provides a scope broader than a single page but narrower than a session. Instead of using @SessionScoped for temporary workflow data, the state exists only for the workflow's duration. This article covers the Faces flow: how pages are connected, how state persists between them, and how entering and exiting the flow manages the Java bean’s lifecycle. The full application also includes MongoDB persistence, dashboard services, preview calculations, and Jakarta Batch processing. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. Here, the emphasis stays on the user interaction represented by Configure → Preview → Review → Exit. Conclusion @FlowScoped provides Jakarta Faces with an effective way to model multi-step user experiences as a single business conversation. Rather than storing temporary workflow state in @SessionScoped or reconstructing it between views, the flow maintains state only while the user progresses through the defined steps and releases it when the flow concludes. For Java-focused enterprise teams, this approach simplifies wizards, onboarding, approvals, checkout flows, and configuration processes by keeping navigation, state, and lifecycle consistent with the user experience.
For the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening. The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified. This is a significant architectural shift. The emerging pattern looks less like: Plain Text prompt → LLM → source code → trust it (the infamous vibe ding pattern) and increasingly like: Plain Text intent ↓ LLM ↓ restricted set of operations ↓ deterministic tools and validation ↓ result Several recent developments in mainstream developer tooling point in exactly this direction. Permission Is Becoming Separate From Intelligence On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations. Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals. That distinction is more profound than it may initially appear. The question is no longer merely: Can the agent perform this operation? It is: Is this agent authorized to perform this operation in this environment? Capability and authority are different things. A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so. A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities. Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering. The Model Is Becoming Replaceable Another development makes the separation even clearer. GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation. This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go. Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear. Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is: Plain Text stable environment stable tools stable constraints stable validation ↑ interchangeable models stable environment stable tools stable constraints stable validation ↑ interchangeable models The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system. Agents Are Separating From Editors VS Code's Agent Host architecture points in another related direction. The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts. Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code: Plain Text human ↓ editor ↓ language server ↓ compiler Agentic development introduces another participant: Plain Text human intent ↓ agent ↓ semantic tools ↓ compiler / tests / environment The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important. A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things. Verification Is Moving From Opinion to Execution A fourth development may ultimately be the most important. GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct. This should sound obvious. We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense. A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system. A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes: Plain Text generate ↓ compile ↓ test ↓ inspect diagnostics ↓ repair ↓ repeat That is substantially more robust than asking a model to inspect its own output and say whether it looks right. The role of the LLM is reasoning. The role of deterministic software remains enforcement. The Interesting Convergence These developments come from different parts of the development stack, but they point toward the same decomposition. An AI programming environment increasingly contains at least four distinct elements: A model that reasons and generatesA vocabulary of operations available to itA capability policy defining which operations it may useDeterministic mechanisms that decide whether the result is valid None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design. What If We Put the Boundary Into the Language? Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted. There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes? This is the idea I have been exploring with an open-source project called BUBAS. BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment. There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary. An order-processing application could, for example, expose operations such as: Plain Text LOAD_ORDER ORDER_TOTAL CUSTOMER_RISK APPROVE REJECT REQUEST_APPROVAL An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed. If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that. Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem. This Is Not a Sandbox Make the distinction carefully. A restricted language does not magically make its host application safe. If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong. The useful guarantee is narrower: Generated business logic can only name operations that the application deliberately made part of its vocabulary. That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment. These operate at different semantic levels. Domain Capabilities Are More Interesting Than Operating-System Capabilities Operating-system permissions are necessary, but business applications eventually need a richer vocabulary. Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful. It still does not answer questions such as: May this program approve an order?May it request approval but not approve directly?Can it read customer risk information?Can it initiate a payment?Can it calculate a premium but not change the underlying policy? Those are domain capabilities. General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction. A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations. We May Be Seeing the New Shape of the AI Programming Stack None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software. But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change. The industry is gradually replacing: Trust the model. with: Give the model precisely defined capabilities and verify what it produces. That is a much more promising engineering principle. BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.
It has been such a joy getting to know our DZone contributors beyond reading their incredible articles. And, pretends to be shocked, developers do, in fact, have lives outside of working on their computers. From building open-source projects to spending quality time with their families, DZone’s community is full of not only subject-matter experts, but passionate individuals with interesting stories to tell. One of those individuals happens to be our Member Spotlight of the week. Mayowa Fajobi may be a newer face on DZone, but he’s already an established leader in tech spaces like platform engineering, AI-driven solutions, and open source strategy. But don’t just take my word for it! Learn more about Mayowa, his background, and the career journey that brought him to where he is today. 1. What first got you interested in technology? I’ve always been fascinated by the idea that technology can turn complex problems into something simple and repeatable. My interest really took shape when I started working with Linux and infrastructure and realised that I could automate processes that would otherwise require hours of manual work. That eventually led me into DevOps, cloud-native technologies and Kubernetes, and then into open source, where I could contribute to the tools I was using rather than simply consuming them. 2. If you could only keep three tools in your developer toolkit, what would they be and why? Linux, Git, and Kubernetes. Linux gives me the foundation for understanding what is actually happening beneath the applications and platforms I build. Git is indispensable for collaboration, experimentation, and open-source contribution. Kubernetes brings the two together at scale and has become one of the most interesting platforms for solving distributed infrastructure problems. If I had to pick a fourth, I’d probably struggle not to say a terminal! 3. What advice would you give someone just getting started in your field? Don’t wait until you feel like an expert before you start building. Learn the fundamentals, build things that break, investigate why they broke, and repeat the process. I would also encourage people to contribute to open source early; even documentation fixes, tests, or small bug fixes can teach you how real-world software is designed and maintained. Most importantly, focus on solving problems rather than collecting technologies. 4. As someone working across platform engineering, AI, and open source, what change do you think will have the biggest impact on how engineering teams build and operate software over the next few years, and what do you think technical leaders should be doing now to prepare for it? I think the biggest shift will be the move from software built around long-running services to software built around intelligent, autonomous, and ephemeral workloads. AI is accelerating this transition. Instead of predictable compute shaped by static deployments, we’re entering a world where agents appear, execute work, pause, resume, and scale dynamically based on context. This will force infrastructure to become far more adaptive, capable of allocating, reusing, and governing compute in real time rather than through fixed-capacity models. Because of this, platform engineering must evolve. It will shift from simply providing clusters and pipelines to providing a unified abstraction layer that manages workload lifecycle, security, observability, and cost for increasingly dynamic systems. Technical leaders should prepare now by: Investing in strong platform foundations and treating infrastructure as a product.Designing for automation, portability, and policy-driven governance rather than rigid setups.Creating safe pathways for engineers to experiment with AI while maintaining strict security and visibility boundaries. Finally, open source will be essential. No single organization can solve these challenges alone. The teams that adapt well will be those that combine solid engineering fundamentals with active participation in the wider ecosystem. 5. What’s your ideal way to spend a weekend? A good balance of technology, music, and time away from technology. I enjoy spending time with family, exploring somewhere new, and playing the piano. I also try to switch off for a while, although I’ll inevitably end up opening my laptop at some point, usually because an interesting open-source problem has been sitting in the back of my mind all week! Check out Mayowa's content here.
Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.
Production bugs frequently arrive with too little evidence. A screenshot captures the final visual state, a crash report identifies a failing stack, and a support ticket describes what appeared to happen. None reliably explains the sequence that produced the failure. Modern applications are asynchronous systems driven by navigation, network responses, feature flags, background work, local persistence, and changing UI state. A useful production bug report therefore needs more than the final frame. It needs a bounded, privacy-safe execution history that reconstructs the path leading to failure. The foundation of a replayable bug report is a semantic event stream. Continuous video recording is expensive, difficult to search, and likely to capture information unrelated to diagnosis. Structured events are smaller and describe meaningful transitions directly. Navigation changes, button actions, state mutations, network outcomes, feature flag evaluations, lifecycle transitions, and persistence failures can use a common event model. Java recordEvent( "ui.action", "checkout.submit", Map.of("screen", "Checkout", "cartState", "ready") ); The event records behavior rather than pixels. Stable identifiers such as checkout.submit are preferable to coordinates or view hierarchy paths because layouts change between releases. Each event should also contain a session identifier, application version, timestamp, and monotonic sequence number. Those fields transform disconnected observations into an ordered execution history. Continuous recording immediately creates a resource constraint. Retaining every event for an entire session eventually consumes excessive memory or storage. A bounded ring buffer solves this by retaining only the most recent diagnostic window. New events replace the oldest once the configured limit is reached. Java private final Deque<ReplayEvent> events = new ArrayDeque<>(); private static final int LIMIT = 500; public synchronized void record(ReplayEvent event) { if (events.size() == LIMIT) events.removeFirst(); events.addLast(event); } Production implementations can enforce both event count and byte size limits because payload sizes vary. High-frequency signals also require control. Scroll offsets, animation callbacks, connectivity heartbeats, and repeated state notifications can overwhelm useful evidence. Sampling or coalescing these signals keeps the buffer focused on transitions that materially affect application behavior. Ordering deserves separate treatment because wall-clock timestamps alone are unreliable. Device clocks can change, concurrent callbacks can receive identical timestamps, and asynchronous tasks can finish in an order different from their creation order. Assigning an atomic sequence number when each event enters the recorder establishes deterministic local ordering. Wall-clock time remains valuable for correlation with server logs, while the sequence number determines authoritative ordering inside the client session. Network operations are especially valuable during reconstruction, but complete request and response bodies usually are not. Recording the HTTP method, normalized route, status code, duration, retry count, failure category, and correlation identifier provides useful evidence without copying sensitive payloads. Java recordNetwork( "POST", "/checkout", statusCode, elapsedMillis, correlationId ); The correlation identifier connects client replay with backend observability. When trace context propagates through an API gateway and downstream services, the same failed interaction can be followed beyond the device. OpenTelemetry context propagation supports this model through trace and span context carried across execution boundaries. A replay system becomes substantially more useful when a client event can lead directly to the corresponding backend trace instead of creating an isolated client-side observability system. Events alone may still be insufficient when identical actions behave differently under different runtime conditions. Selective state snapshots fill that gap. A snapshot should not serialize the complete application object graph. It should capture small diagnostic facts that influence behavior, such as authentication status, active feature flags, connectivity mode, pending operation counts, cache generation, lifecycle state, and other safe application state. Java recordSnapshot(Map.of( "screen", "Checkout", "network", "cellular", "paymentState", "submitting", "feature.checkoutV2", true )); Snapshots can be generated at meaningful boundaries such as screen entry, transaction start, backgrounding, synchronization completion, or error detection. During reconstruction, they explain conditions surrounding an event without attempting to reproduce every byte of runtime memory. This keeps diagnostic bundles small while preserving state likely to affect execution. Privacy must be enforced during capture rather than treated as an upload-time cleanup operation. Session replay platforms commonly provide client-side masking because sensitive values should not leave the application in their original form. The same principle applies to a custom recorder. Event attributes should follow explicit allowlists, while passwords, authorization headers, tokens, payment information, message contents, email addresses, and unrestricted text fields should be excluded by default. Java private String sanitize(String key, String value) { if (SENSITIVE_KEYS.contains(key)) return "[REDACTED]"; return value; } Key filtering alone is insufficient because sensitive values can appear under unexpected field names. Stronger implementations can combine schema allowlists, endpoint-specific policies, value-pattern detection, maximum lengths, and explicit data classifications. Redaction should occur before information enters the ring buffer. Once sensitive data has reached memory, disk, crash attachments, or telemetry queues, later sanitization becomes considerably harder to guarantee. The recorder must also survive the failure being diagnosed. An entirely in-memory history disappears during process termination, watchdog kills, native crashes, or operating system eviction. Periodic checkpointing to a small protected file allows the next launch to recover the tail of the previous session. Writes should remain asynchronous and bounded so diagnostic instrumentation does not introduce latency or instability into normal execution. Crash time behavior should remain minimal. Attempting complex serialization after a fatal condition can itself fail. A safer design periodically persists compact checkpoints during healthy execution, and treats crash handling as a final marker whenever possible. On the next launch, the previous checkpoint, crash metadata, application version, device characteristics, and correlation identifiers can be assembled into a diagnostic bundle. Upload requires failure handling as well. A diagnostic system cannot assume connectivity exists immediately after a crash or restart. Bundles can enter a small persistent queue and upload when network conditions permit. Successful delivery removes the local copy, while repeated failures follow bounded retry and retention policies. This prevents diagnostic infrastructure from becoming another source of uncontrolled storage consumption or retry storms. Replay does not need to mean pixel-perfect visual reproduction. For engineering diagnosis, a deterministic event timeline is often more useful. Plain Text 10:42:11.018 screen.enter Checkout 10:42:14.201 ui.action checkout.submit 10:42:14.233 network.start POST /checkout 10:42:15.107 network.finish status=200 10:42:15.116 persistence.error order_write_failed 10:42:15.120 ui.state payment=failed A screenshot from this incident shows only a failed checkout screen. The timeline reveals something fundamentally different, as the server operation succeeded, but local persistence failed afterward. That distinction changes the investigation immediately. Instead of examining API availability or retry behavior, diagnosis can move directly toward local storage, transaction handling, or post-response state transitions. Structured replay can extend beyond manual inspection. Development builds can consume sanitized event sequences to configure test doubles, restore relevant feature flags, reproduce network outcomes, and drive selected state transitions. Full determinism is not always possible because operating system scheduling, race conditions, third-party services, and timing-sensitive behavior introduce variability. Even without executable replay, a causal timeline drastically reduces the search space by preserving conditions that conventional logs frequently lose. The recorder should remain diagnostic infrastructure rather than becoming another analytics pipeline. Capturing every possible signal increases CPU usage, storage requirements, privacy exposure, and noise. A small event vocabulary is usually more effective, just like navigation transitions, meaningful interactions, network starts and outcomes, state changes, persistence operations, lifecycle events, and failures. Instrumentation quality matters more than event volume. A practical rollout can begin around workflows responsible for the most expensive production failures rather than instrumenting every screen. Checkout, authentication, synchronization, uploads, and other difficult flows can receive semantic events and snapshots first. When an incident demonstrates that an important transition is missing, the event model can evolve deliberately. This keeps the recorder maintainable and prevents instrumentation from becoming an uncontrolled collection of arbitrary log statements. Screenshots remain useful evidence, but they should represent one frame inside a richer diagnostic record rather than the entire debugging strategy. A bounded semantic event buffer, deterministic ordering, selective state snapshots, trace correlation, capture time redaction, crash-safe persistence, and resilient upload can transform vague production reports into reconstructable execution histories. The result is more than improved logging. It is a practical application flight recorder that preserves the context normally lost at the exact moment a difficult production failure occurs.
Every production incident starts with a simple question: "Has this happened before?" I've lost count of how many incident bridges I've joined where that question came up within the first few minutes. Before anyone proposes restarting a service or rolling back a deployment, someone inevitably starts searching. They look through Slack conversations from previous outages, browse old postmortems, compare dashboards with similar incidents, or dig through runbooks to see whether another team has already solved the same problem. Do you notice what's happening here? The engineers aren't trying to demonstrate how much they know about distributed systems. They're trying to remember. That observation has become increasingly important as AI assistants find their way into engineering organizations. Today's large language models are remarkably good at explaining Kubernetes concepts, debugging stack traces, writing SQL queries, or summarizing log files. Those capabilities are valuable, but they only solve part of the problem. During an incident, reasoning is rarely the bottleneck. Finding the right context is. The engineer who resolves an outage the fastest isn't always the one with the deepest theoretical knowledge. More often, it's the engineer who remembers that a similar issue occurred eight months ago after a database failover, or who knows that a particular service has historically exhibited the same failure pattern after specific deployment changes. Experience is a form of memory. If we want AI to become a trusted operational partner instead of just another chatbot, we need to think about memory as carefully as we think about intelligence. Intelligence Answers Questions. Memory Solves Problems. Large language models are exceptional at answering questions because they've been trained on an enormous amount of public knowledge. Ask an LLM to explain consensus algorithms, Kubernetes scheduling, or distributed tracing, and you'll probably receive a detailed, technically accurate explanation within seconds. Production incidents, however, ask very different questions. Instead of asking: What is Kubernetes? Engineers ask: Why did our Kubernetes cluster start failing after yesterday's deployment?Has this service failed in the same way before?Which runbook actually worked the last time?Who owns this dependency today?What changed in the last hour that could explain this behavior? Those aren't questions about computer science. They're questions about organizational memory. The answers don't exist in a foundation model's training data because they're unique to every organization. They live in deployment histories, incident timelines, internal documentation, architecture decisions, monitoring dashboards, chat conversations, and postmortems accumulated over years of operating software. An AI that understands distributed systems but lacks access to this operational history is like an experienced consultant joining your incident bridge for the first time. It may offer useful suggestions, but it doesn't know your environment, your systems, or your team's accumulated experience. That's why intelligence alone isn't enough. Every Incident Is a Search Problem One pattern I've noticed is that incident response often looks less like debugging and more like information retrieval. Think about what engineers actually do during the first fifteen minutes of a major incident. One person opens dashboards to identify where the failure started. Another compares the current deployment with the previous version. Someone searches Slack for keywords that resemble the current symptoms. Another engineer opens the last postmortem for the affected service. Meanwhile, the incident commander tries to understand which teams need to be involved. None of these activities involve writing complex algorithms. They're all attempts to reconstruct context. If you mapped the engineer's workflow, it might look something like this: Alert → Metrics → Logs → Deployment History → Previous Incidents → Runbooks → Slack Discussions → Architecture Documentation → Decision The common thread is that engineers are constantly retrieving information before making decisions. That retrieval process is exactly where AI can provide the most value—not by replacing engineering judgment, but by dramatically reducing the time required to gather relevant context. Not All Memory Is the Same When we talk about memory in AI, it's easy to think only about conversation history or a vector database. In practice, incident response depends on several different kinds of memory, each answering a different set of questions. Incident memory This is the collective history of operational failures. Previous incidents, timelines, root causes, postmortems, and lessons learned all fall into this category. During an outage, one of the first questions engineers ask is whether they've seen the problem before. An AI that can retrieve similar incidents and explain how they were resolved immediately provides value because it shortens the investigation. Operational memory Runbooks, playbooks, escalation procedures, and service ownership represent another form of memory. These artifacts capture how an organization expects engineers to respond under different circumstances. Instead of generating a generic remediation plan, an AI can recommend the procedure that has already been validated by the organization. Infrastructure memory Production systems change constantly. Deployments, feature flags, infrastructure updates, configuration changes, and dependency upgrades all influence system behavior. Understanding what changed recently is often more useful than understanding how a technology works in theory. Organizational memory Some of the most valuable operational knowledge never reaches formal documentation. Engineers discuss recurring issues in Slack, record architectural decisions in design documents, and exchange troubleshooting tips during retrospectives. Over time, this becomes institutional knowledge that experienced engineers rely on instinctively. AI should be able to surface that knowledge instead of forcing every engineer to rediscover it. Memory Changes the Quality of Recommendations Imagine two AI assistants responding to the same latency alert. The first assistant says: CPU utilization is high. Consider restarting the service. It's not necessarily wrong, but it's also not particularly helpful. Now imagine a second assistant with access to organizational memory says: A similar incident occurred three months ago after deployment version 6.4. During that incident, restarting the service temporarily reduced latency, but the underlying cause was an inefficient database query introduced by the deployment. The query was reverted, and latency returned to normal within six minutes. A deployment with similar changes occurred eighteen minutes before the current alert. I recommend validating query performance before restarting the service. Neither assistant is more intelligent in the traditional sense. The second assistant is simply making better use of memory. That additional context changes the recommendation from a generic suggestion into operational guidance grounded in the organization's own experience. Building Memory Into AI Systems Memory isn't a single database or a single technology. It's an architectural capability that combines multiple sources of operational knowledge into a coherent context for reasoning. A production-ready incident assistant might continuously ingest information from observability platforms, deployment pipelines, service catalogs, incident management systems, internal documentation, and communication channels. Rather than asking engineers to manually gather information from each source, the AI assembles the relevant context before generating a recommendation. The language model is still responsible for reasoning, summarization, and communication. The memory layer ensures that reasoning is grounded in facts that are specific to the organization rather than generic patterns learned during training. In many ways, this mirrors how experienced engineers work. They don't solve incidents by relying only on theoretical knowledge. They combine technical understanding with years of accumulated operational experience. That's exactly the capability our AI systems should emulate. Final Thoughts As large language models continue to improve, it's tempting to believe that more intelligence alone will solve the challenges of operational AI. My experience suggests otherwise. The most effective incident response systems aren't necessarily the ones with the largest models or the most sophisticated prompts. They're the ones that help engineers remember. They surface the right runbook, identify the last time a service failed in the same way, highlight the deployment that introduced the problem, and connect today's symptoms with yesterday's lessons. In other words, they make organizational experience accessible when it's needed most. Incident response has always been a combination of reasoning and memory. AI has made remarkable progress on the first half of that equation. The next step isn't simply building smarter models—it's building systems that remember.
An AI agent rarely follows a long plan exactly as first proposed. Tool results expose missing facts, external systems change, policies arrive through human input, and a model may discover that an earlier assumption was wrong. ReAct-style agents were explicitly designed around this interleaving of reasoning, action, observation, and plan updates rather than one immutable plan. The engineering difficulty begins when those actions are durable side effects. In Temporal, a completed Activity is not an intention that can be edited out of a revised plan; its completion and result are part of the Workflow Event History. Replanning therefore has to reconcile a new decision with an already-real past. A Plan Is Intent; History Is Fact The cleanest mental model separates plan state from committed execution state. The plan is provisional: it represents what the agent currently intends to do. Committed state represents facts established by completed Activities, accepted external messages, and other durable events. Temporal stores Workflow history as an append-only sequence and uses that history to reconstruct Workflow state during replay. An ActivityTaskCompleted event contains the serialized Activity result, so a later model decision cannot make that completion disappear. That distinction changes the shape of the agent loop. Instead of asking an LLM to generate a complete plan and then blindly executing it, each replanning call should receive the goal, the latest observations, the prior plan revision, and a normalized set of committed effects. The model may abandon remaining steps, but completed steps should enter the next prompt as constraints on the current world state. Java public AgentResult run(Goal goal) { while (!goalSatisfied(goal, committed)) { Plan plan = planner.replan( new PlanningContext(goal, revision, committed)); for (Action action : plan.readyActions()) { if (alreadySatisfied(action, committed)) { continue; } ActionReceipt receipt = tools.execute(action, action.operationId()); committed.add(receipt); if (receipt.requiresReplan()) { break; } } revision++; } return summarize(committed); } In this shape, planner.replan(...) and tools.execute(...) are Activities, not arbitrary network calls from Workflow code. That matters because Workflow code must remain replay-safe, while non-deterministic operations such as LLM invocations belong behind durable boundaries. Temporal records Activity results in history and returns those recorded results during replay instead of re-running the external call as part of Workflow replay. Temporal’s Java AI integration applies the same principle by executing model calls and external tools through Activities, while deterministic operations may remain inside the Workflow. Replanning Should Consume Effects, Not Merely Step Status A boolean such as completed=true is usually too weak for reconciliation. A completed tool call should return an execution receipt describing the externally relevant postcondition: resource identifiers, amounts, versions, timestamps when materially relevant, and whether a compensating operation exists. The agent then reasons from effects rather than from labels such as “step three succeeded.” Consider a reservation agent that initially plans to reserve inventory, create a shipment, and notify a customer. Inventory reservation succeeds, but shipment creation reports a destination restriction. A revised plan must not schedule another reservation merely because the original plan was discarded. The durable input to replanning should state that a specific reservation already exists and identify it. The next plan might change the shipping method, release the reservation, or escalate for approval, but it should not pretend that the reservation never happened. A second guard is a planning epoch. External input can arrive while a planning Activity is outstanding; Temporal schedules Workflow Tasks when Signals or Updates arrive, and those messages can mutate durable Workflow state. A plan generated from revision 12 should therefore not be committed blindly if the state has advanced to revision 13 before the planner returns. The planner result can carry the revision used as its input, and the Workflow can discard stale output and request a fresh plan. This is optimistic concurrency control applied to agent reasoning rather than database rows. Retries make stable identity equally important. Temporal recommends idempotent Activities because Activities can be retried, and its documentation specifically notes that idempotency keys are appropriate for critical side effects. A plan revision must therefore not double as an operation identity. A reservation that remains semantically the same across replans should retain the same business operation key; a genuinely new reservation should receive a new key. Java public ReservationReceipt reserve(ReservationCommand command) { return inventory.reserve( command.sku(), command.quantity(), command.operationId()); } Here, operationId is passed through to a downstream API that enforces idempotency. The receipt returned to the Workflow becomes durable evidence of the reservation. This also prevents a subtle failure mode in which a model repeats a tool call after losing conversational context even though the Workflow already possesses proof of the earlier effect. Cancellation and Compensation Are Forward Actions A changed plan often creates pressure to “cancel the old step,” but cancellation has a precise boundary. Cancellation can stop work that is still in flight; it cannot retroactively cancel an Activity that has already completed. For long-running Activities, Temporal delivers cancellation cooperatively through Activity heartbeats, so heartbeat configuration determines how promptly an Activity can observe a cancellation request. Once an effect has committed, reconciliation becomes a domain operation. If the effect is reversible, compensation is normally the correct mechanism. Temporal documents the Saga pattern as a sequence of local operations paired with compensating actions, typically executed in reverse order when later work makes earlier effects undesirable. The important semantic point is that compensation creates new history. It does not rewrite old history. A released reservation follows a created reservation; a refund follows a charge; a revocation follows a grant. Java private void reconcile(Plan revisedPlan) { for (ActionReceipt receipt : compensationsInReverseOrder(revisedPlan, committed)) { ActionReceipt reversal = tools.compensate(receipt); committed.add(reversal); } } Not every action has a true inverse. An email cannot be unsent, an external party may already have observed a published event, and a physical operation may have crossed an irreversible boundary. Such effects should be modeled as facts that constrain future planning, not as failures of rollback. The revised plan can issue a correction, create a follow-up notification, or route the case to a human decision, but the historical effect remains part of the state presented to the agent. Durable Agents Need Forward-Only Semantics Replanning also needs to stay distinct from Workflow code versioning. A model changing its runtime plan is ordinary application behavior: a new planning Activity produces a new durable decision after new evidence arrives. Changing deployed Workflow code is different because replay must still produce commands compatible with existing Event History; Temporal provides versioning and patching mechanisms for those code changes. Mixing the two concepts leads to brittle systems in which model variability is handled as deployment variability or, worse, non-deterministic Workflow logic. External corrections fit the same forward-only model. Temporal Signals and Updates can change running Workflow state, and accepted messages become durable inputs that can trigger another planning turn. For very long-running agents, Continue-As-New creates a fresh Event History while carrying forward explicit application state, which makes the committed-effect ledger an important part of the continuation payload rather than transient model memory. The central design rule is simple: an agent may revise intentions at any time, but durable execution never revises facts. Completed Activities should be represented as committed effects with stable identities and useful receipts; in-flight work may be canceled cooperatively; reversible effects may be compensated; irreversible effects must constrain the next decision. Temporal’s history then becomes more than a recovery mechanism. It becomes the authoritative boundary between what the agent merely planned and what the surrounding world has already observed. That boundary allows adaptive AI behavior without sacrificing replay safety, idempotency, auditability, or operational correctness.
If your team distributes AI development skills (as a Claude Code or Cursor plugin or a shared rules file), you own a catalog. That catalog covers things like: How to structure a serviceWhat must pass before a commitWhich internal library to use instead of rolling your own Skills load cheaply, thanks to progressive disclosure. Only the name and one-line description sit in context until the description matches the work. So when token spend climbs, the intuitive read is context bloat, and the obvious lever is tuning descriptions to fire less. That work is worthwhile as it improves routing quality. But it addresses only one of the two ways a skill costs money. The other is procedural amplification: a skill that encodes a fixed procedure as prose, so the model reconstructs it, deliberating, calling tools, checking results, on every invocation. That doesn’t show up as a large context payload. It shows up as extra inference round trips, and an aggregate cost dashboard cannot see it. Predictability does not prove a step should be automated. It identifies places where you should ask whether you’re paying a model to make a decision the system has effectively already made. A few terms come up more than once below, defined here so they don’t slow you down later: TermMeaningSkillA packaged set of instructions Claude Code loads only when the task matches it.OTelShort for OpenTelemetry, the open standard Claude Code uses to report what it did and how much it cost.API request / round tripOne call to the model and its response. This article counts these to measure cost.HookA small script Claude Code runs automatically at a specific moment, such as right after every tool call.p50 / p90 / p95 / p99Percentile timing. p50 is the typical (median) call; p99 is close to the worst case you’ll see.AsyncA setting that lets a hook run in the background instead of making the agent wait for it to finish.NDJSONOne JSON record per line in a file. Simple to append to and to stream. Use the Vendor’s Telemetry With Small Customization I’ll correct a claim I believed myself and have seen repeated: OpenTelemetry only provides aggregate counters, so you have to build your own event pipeline. That’s false, and starting from scratch costs you a week. Claude Code’s OTel export includes events, and they’re richer than I expected: EventCarriesclaude_code.api_requestskill.name, cost_usd, input_tokens, output_tokens, duration_ms,event.sequence, prompt.idclaude_code.tool_resulttool_use_id, tool_name, success, duration_ms skill.name on the API request is the important one: inference round trips are already attributed to the skill that was active. cost_usd is documented as an estimate, not billing data. The One Field to Add Getting command detail out of the native exporter requires OTEL_LOG_TOOL_DETAILS=1, which attaches full_command, bash_command, and file_path to tool events. That’s exactly the payload you cannot centralize off a fleet of developer laptops: branch names, commit messages, customer identifiers, the occasional pasted token. So a small local hook fills one gap: a privacy-preserving shape of the command, joined to native telemetry on tool_use_id, which the docs describe as matching the tool_use_id passed to hooks, allowing correlation between OTel events and hook-captured data. Plain Text native OTel ──┐ │ api_request: skill.name, cost, tokens, round trips ├── join on tool_use_id ──▶ analysis │ tool_result: tool_use_id, tool_name, success local hook ──┘ The hook doesn’t rebuild the trace. It emits the join key plus the one thing native telemetry can’t safely give you. Anything OTel already reports (success, duration_ms, skill.name, token counts) is deliberately not duplicated. Two sources of truth for one field will eventually disagree, and you’ll trust the wrong one. The command-to-shape transform itself is a per-binary allowlist, not a secret detector, and it has a sharp edge worth knowing: The transform: git commit -m "fix auth for acme" becomes git commit -m <ARG>.The edge case: a value-taking flag like --token abc123 leaks its argument as a bare token unless you explicitly track which flags consume the next word. Get the allowlist wrong, and the safety argument for the whole pipeline goes with it. What the Hook Costs (Measured) Tool hooks run on the critical path, so I benchmarked one: 1,000 synthetic PostToolUse payloads through the real script, mixed across Bash/Read/Edit/Write/Grep/Skill. Plain Text p50 p90 p95 p99 max mean subprocess 24.28ms 26.91ms 27.91ms 32.31ms 52.42ms 24.93ms in-process 0.08ms 0.15ms 0.16ms 0.20ms 0.35ms 0.10ms macOS, arm64, 10 cores, Python 3.9.6, n=1000 after 25 discarded warmups. The interesting part isn’t the total; it’s the split. About 24.2 ms of that 24.3 ms p50 is Python interpreter startup. The actual work — scrubbing, serializing, appending — costs 0.08 ms. That points the fix somewhere I would not have guessed: Don’t optimize the scrubber. It’s already three orders of magnitude below the launch cost.Do fix how the hook is launched. Register it async: true so nobody waits, or make it a thin client to a long-lived local collector so you pay startup once per session instead of per tool call. Had I estimated instead of measured, I’d have guessed low single-digit milliseconds and gone tuning the scrubbing code. Wrong target entirely, which is the argument for measuring rather than estimating, in miniature. Event records came in at 364 bytes each. One engineer at roughly 40 sessions a week and 60 tool calls a session generates about 2,400 records, under 1 MB of raw NDJSON weekly. It stays small-data at any plausible team size, because the enrichment record only carries what native OTel doesn’t. Measuring Your Own Catalog Skills have two token costs that behave differently, and knowing which one dominates for you tells you where to look first: Cost componentPaid whenScales withDescriptionEvery request, every session, whether the skill fires or notCatalog sizeBody (procedural amplification)On invocation, then re-sent on each subsequent request in the sessionHow often the skill fires, and how long sessions run Measuring the 25 skills in Claude Code’s official plugin marketplace, a catalog anyone can install and re-measure, gives: Plain Text body chars: min 989 median 11,395 max 32,625 (33x spread) desc chars: min 105 median 357 max 906 always-resident total (all 25 descriptions): 9,703 chars Read those two numbers against each other. Carrying every description in the catalog costs less than a third of one large body. The single biggest skill is roughly 3.4 times the entire always-resident footprint, and it’s charged again on every request for the rest of any session that invokes it. That’s the shape that makes procedural amplification worth hunting. If your always-resident total instead dwarfs your median body, the lever really is catalog size and description tuning. You can stop there. (These are characters, not tokens. The ratio shifts with how much code versus prose a body contains. Count with the real tokenizer, messages.count_tokens, which is free and rate-limited only by requests per minute; don’t reuse a chars/4 estimate or another vendor’s tokenizer, and don’t reuse an old Claude count either: the tokenizer used by Claude 4.7 and later yields roughly 30% more tokens for the same text than earlier models.) Why This Is Worth Measuring at All The subject isn’t token optimization. It’s finding misplaced probabilistic computation: places where the system already knows the next operation and is paying a model to rediscover it. If the next operation is…It belongs in…Genuinely uncertainA skill, policy and judgment, model reasoningAlready determinedA script, deterministic execution A skill that spells out a fixed procedure in prose has put that boundary in the wrong place. The industry is converging on the same line from the runtime side: programmatic tool calling exists precisely to keep deterministic multi-tool sequences out of the inference loop. Key Takeaways Aggregate cost accounting answers the question you already knew to ask. Event-level behavioral traces let you find the one you didn’t, and in Claude Code, most of that stream already exists, attributed to the skill, waiting on one privacy-preserving field to become useful. Procedural amplification leaves no trace in a token counter, raises no error, and turns no dashboard red. It’s visible in the order of the calls, and in how many round trips it takes to get through them. Measured, not estimated: Hook overhead is about 24 ms per call, almost entirely interpreter startup, not the redaction logic itself.Measured, not estimated: In one real catalog, a single skill body costs 3.4× what the entire always-resident description set costs.Next step: Measure by the actual stretch of work, not by an arbitrary number of calls. Group everything that happens while one skill is active, however long that turns out to be, rather than picking a fixed count upfront. When one model turn fires off several tool calls at once, count that as the single decision it was, not several. And don't trim the long, expensive stretches out of the data before analyzing it. Those are usually the exact cases worth finding.
Mobile networks are notoriously unreliable. A common scenario is when a user taps "Pay Now" in an app: the payment request reaches the server and is processed, but the network response never reaches the phone. The client assumes the request failed and retries, leading to the charge running twice. This is precisely the kind of bug that idempotency solves. Idempotency means that repeating the same operation has no additional effect, and the second attempt should recognize it's a duplicate and do nothing new. In practice, mobile engineers must treat payment or order APIs as idempotent by attaching unique operation identifiers to requests and deduplicating them on the backend. With this approach, even if the network drops a response or the user double-taps a button, the user is charged only once. Achieving this involves coordination between the app and the server. On the client side, every payment or mutation request is given a persistent unique ID. For example, the app might generate a new UUID when the user submits a payment, save that operation in a local "outbox" or queue, and include the ID in the HTTP request: Swift let opID = UUID().uuidString var request = URLRequest(url: URL(string: "/api/payments")!) request.httpMethod = "POST" request.setValue(opID, forHTTPHeaderField: "Idempotency-Key") request.httpBody = /* JSON payload of the payment */ This custom header (Idempotency-Key) carries the operation's identity to the server. (Modern APIs often expect this header to ensure that POST is treated safely.) The client's logic must be persistent; before sending, it writes the operation (ID and payload) into local storage (e.g., SQLite or SharedPreferences) so that it can recover and retry if the app closes or the network is down. This pattern, sometimes called the outbox pattern, means the user's intent is recorded immediately. A background dispatcher can then drain the queue on network availability or app restart; it tries each pending request. If a send fails (timeout, no internet, 5xx error), the entry remains in the queue for a later retry. This ensures at-least-once delivery and the server will eventually see the request, even if the app crashes or the network is flaky. But at-least-once alone would cause duplicates, so the server must be ready. When the backend receives the request with its Idempotency-Key, it first checks a deduplication store (for example, a database table keyed by this ID). Java String key = request.getHeader("Idempotency-Key"); PaymentResponse prev = idempotencyStore.lookup(key); if (prev != null) { // We have processed this request before - return the original response return prev; } // No record of this key; proceed with processing PaymentResponse result = processPayment(request.getBody()); // Store the result before returning it idempotencyStore.insert(key, result); return result; If the key already exists, the server simply returns the stored result without charging again. This ensures that the second (or third) time the client re-sends, the user doesn't get double-charged. Stripe's API, for instance, works exactly this way: it saves the outcome of the first request for a given idempotency key, and any retry with the same key returns the same result. In effect, the combination of at-least-once delivery (the client keeps retrying) plus idempotent handling on the server yields an effectively-once outcome. The server's idempotency store can be implemented with a simple database table that records each key and the operation's result. For instance, a processed_payments table might use the idempotency key as a primary key or unique constraint. The service then does an atomic INSERT ... ON CONFLICT DO NOTHING (PostgreSQL syntax) or equivalent. If the insert succeeds, the code proceeds with the payment and stores the result; if it fails because the key already exists, it knows this is a duplicate and can fetch the prior result. Wrapping the insert and the business operation in one database transaction avoids a race condition, as either both the key and payment record are written, or neither is. In SQL terms: SQL BEGIN; INSERT INTO payments(idempotency_key, user_id, amount) VALUES (:key, :userId, :amount) ON CONFLICT (idempotency_key) DO NOTHING; -- Check how many rows were inserted: IF (INSERT was successful) THEN -- This is the first time seeing this key; perform the payment CALL process_payment(...); -- The payment service may record a transaction ID, etc. COMMIT; ELSE -- Key already existed: rollback any partial work ROLLBACK; -- Retrieve and return the original payment result END IF; Even if two identical requests arrive concurrently, the unique constraint ensures only one succeeds in its insert. The other can detect the conflict and simply return the saved response. The system design sandbox guide describes this approach as "The database enforces uniqueness and no separate check needed. This works well when the idempotency record belongs in the same database as the business data, since you can wrap both in a single transaction." On the mobile side, it's also wise to guard against duplicates before the request is even sent. A simple in-memory or on-disk set of "seen" IDs can help reject retry loops after a crash or double tap. In Swift: Swift final class OperationDeduplicator { private var seen: Set<String> = [] func shouldProcess(_ id: String) -> Bool { return seen.insert(id).inserted } } This OperationDeduplicator returns true only the first time an ID appears. Persisting this set across app launches (for example in Core Data or a file) makes the app resilient to a crash after the payment is sent but before the response arrives. On relaunch, the app knows it already handled that operation and won't enqueue it again. It's important to integrate these pieces smoothly. A typical mobile flow might look something like this: the user submits a payment form, the app immediately generates a new opID (a UUID) and creates an operation record { id: opID, payload: {amount, items, ...} }. This record is saved locally. Then a background task picks it up, attaches opID as the Idempotency-Key header, and sends it. If the network call times out, the record stays queued. When the app regains connectivity or restarts, the dispatcher tries again. Because the same opID is used each time, the server knows to treat all retries as one. Only after the server successfully processes the payment does the app remove the operation from its queue. This pattern ensures retries and crashes do not cause duplicate side effects. Some systems even use more granular controls. For example, if the backend involves multiple microservices, one service might call others, and each service should propagate the same idempotency key or a related correlation ID so that the entire transaction remains idempotent. Distributed tracing can help debug how a request flowed through the system. Ultimately, the goal is to capture the entire user action from UI tap through backend processing and make sure it's only applied once globally. This often means also having the backend return the same HTTP status and response body on every retry, so the client never gets an unexpected error. Developers should simulate network failures and verify that retries do not cause double effects. Most important is to observe real production behavior, as logs or traces with the operation ID can tie multiple client attempts to a single transaction. If everything is correct, the system achieves effectively-once behavior where the payment occurs exactly once no matter how many times the client tries. As systemdesignsandbox summarizes, "at-least-once delivery + idempotent consumer = effective exactly-once". In practice, this means mobile apps can assume failures are not fatal and they can safely retry with the same key, knowing the server will protect against duplicates. Conclusion In summary, retry-safe mobile operations require treating each user action as an idempotent transaction. The client must persist a unique operation key and reuse it across retries, while the server must detect previously processed keys and prevent duplicate side effects. Combining durable client operations with server-side idempotency allows payments and other critical transactions to survive timeouts, crashes, and unreliable networks without being executed twice.
Agile
Career Development
Methodologies
Team Management
Can Your Team Name the Work It Already Runs With AI?
September 25, 2026
by Stefan Wolpers
CORE
One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph
September 25, 2026
by Akhil Madineni
CORE
How to Build an Asynchronous AI-Content Review Workflow in C#
September 23, 2026
by Brian O'Neill
CORE
AI/ML
Big Data
Databases
IoT
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
September 29, 2026
by Jubin Soni, FBCS
CORE
How to Test POST API Requests With Playwright TypeScript
September 29, 2026
by Faisal Khatri
CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
September 29, 2026
by Gaurav Gaur
CORE
Cloud Architecture
Integration
Microservices
Performance
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
September 29, 2026
by Jubin Soni, FBCS
CORE
How to Test POST API Requests With Playwright TypeScript
September 29, 2026
by Faisal Khatri
CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
September 29, 2026
by Otavio Santana
CORE
Frameworks
Java
JavaScript
Languages
Tools
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
September 29, 2026
by Jubin Soni, FBCS
CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
September 29, 2026
by Otavio Santana
CORE
Jakarta Faces Flow Scope: Managing Multi-Step UX Without Session State
September 29, 2026
by Otavio Santana
CORE
Deployment
DevOps and CI/CD
Maintenance
Monitoring and Observability
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
September 29, 2026
by Jubin Soni, FBCS
CORE
How to Test POST API Requests With Playwright TypeScript
September 29, 2026
by Faisal Khatri
CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
September 29, 2026
by Gaurav Gaur
CORE
AI/ML
Java
JavaScript
Open Source
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
September 29, 2026
by Jubin Soni, FBCS
CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
September 29, 2026
by Gaurav Gaur
CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
September 29, 2026
by Otavio Santana
CORE