DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

DZone Spotlight

Wednesday, September 30 View All Articles »
Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus

Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus

By Akhil Madineni DZone Core CORE
Agent systems are increasingly decomposed into specialized services: a planning agent delegates research, a research agent invokes retrieval and synthesis, and a compliance agent validates the result before an action is allowed. Open standards such as A2A formalize communication between independent agents, but transport interoperability is only part of the production problem. Requests can be lost after acceptance, retries can duplicate expensive work, callers can disappear while remote tasks continue, and multi-agent chains can become difficult to reconstruct. Temporal Nexus addresses a different layer. It connects Temporal applications through durable service contracts, allowing an agent capability to behave less like a fragile HTTP handoff and more like a reliable distributed operation. When an Agent Call Becomes a Distributed Transaction Boundary A conventional agent handoff often looks like ordinary RPC: serialize a task, send it to another service, wait for a response, and retry failures. That model works for short, stateless interactions. It becomes fragile when delegated work lasts minutes or hours, crosses ownership boundaries, depends on databases or model providers, or triggers side effects. Retry behavior, deduplication, cancellation, timeout ownership, and result delivery then become part of the application protocol. Nexus moves those concerns into Temporal’s execution model. A Nexus Endpoint routes a request to a target Namespace and Task Queue while hiding those implementation details from the caller. The caller depends on a named service contract rather than the handler’s Workflow type, queue, or deployment topology. Temporal describes the relationship as peer-to-peer: caller and handler Workflows remain independent executions while Nexus provides the durable boundary between them. This distinction matters for agent platforms because a durable agent service should expose a business capability, not internal orchestration mechanics. A “research” operation can remain stable even if its implementation changes from one Workflow to a graph of Activities, model calls, human approval, or additional Nexus calls. Temporal supports multi-level Nexus composition, with each hop represented as a separate durable operation. A Nexus Contract Fits the Agent Capability Boundary In the Java SDK, a Nexus Service defines the callable contract and Nexus Operations define individual capabilities. A compact contract for a research agent can remain deliberately narrow: Java @Service interface ResearchAgentService { @Operation AgentResult research(AgentTask task); } The contract carries only the input and output required by the collaboration boundary. Temporal’s Java guidance recommends service APIs containing operation names plus serializable input and output types; JSON and Protobuf are practical choices when multiple SDK languages are involved. The caller therefore does not need access to the handler’s prompts, memory representation, tool graph, or Workflow implementation. Long-running agent work should normally use an asynchronous Nexus Operation backed by a Workflow. Temporal reserves synchronous operations for reliable, predictably low-latency paths that finish within the ten-second handler deadline. Asynchronous operations can represent long-running work and return an operation token for tracking completion without holding a conventional request open. In Temporal Cloud, the maximum Nexus Schedule-to-Close timeout is 60 days. A handler can expose an agent Workflow without introducing an HTTP controller or polling API: Java @OperationImpl public OperationHandler<AgentTask, AgentResult> research() { return WorkflowRunOperation.fromWorkflowMethod( (ctx, details, task) -> Nexus.getOperationContext() .getWorkflowClient() .newWorkflowStub( ResearchAgentWorkflow.class, WorkflowOptions.newBuilder() .setWorkflowId("research-" + details.getRequestId()) .build()) ::run); } WorkflowRunOperation.fromWorkflowMethod maps the Nexus Operation to a Workflow execution. Temporal documents the Nexus request ID as stable across retries, making it suitable for constructing a deduplication-oriented Workflow ID when no stronger business identifier exists. A business identifier is generally preferable because it aligns duplicate suppression and operational correlation with domain semantics. Durability Changes the Meaning of Retry and Cancellation The most important Nexus behavior is execution semantics. Once a caller Workflow schedules an operation, Temporal atomically hands the command to Nexus machinery. Delivery uses at-least-once execution, automatic retry, rate limiting, concurrency limiting, load balancing, and circuit breaking. A handler can therefore be invoked more than once for the same operation, which makes idempotency essential for external side effects. Temporal also documents stronger duplicate protection by backing the operation with a Workflow whose ID reuse policy rejects duplicates. That behavior is especially relevant to agents because model inference may be nondeterministic and repeated tool calls can duplicate actions such as ticket creation, notifications, or database mutations. Durable execution does not remove the need for idempotency keys at external boundaries. Instead, it provides a stable execution context in which prior decisions and completed steps can be recorded and replayed consistently. Cancellation also becomes a protocol-level feature rather than a best-effort HTTP convention. Canceling a caller Workflow propagates cancellation to pending Nexus Operations and their backing handler Workflows. Termination is different: terminating the caller abandons pending operations and does not send cancellation to the handler Namespace, so remote work can continue. Graceful cancellation is therefore the safer control path for agent chains that may need cleanup or compensation. Timeouts express separate service-level expectations. Schedule-to-Start bounds how long an operation may wait to begin, Start-to-Close bounds asynchronous execution after start, and Schedule-to-Close bounds the complete lifecycle. Nexus retries retryable failures within those limits, while non-retryable failures resolve the operation and become visible to the caller. The Caller Stays Simple While the Runtime Carries State From a caller Workflow, the remote agent looks like a typed service stub. Retry loops, callback plumbing, and result polling do not need to appear in business code: Java private final ResearchAgentService research = Workflow.newNexusServiceStub( ResearchAgentService.class, NexusServiceOptions.newBuilder() .setOperationOptions( NexusOperationOptions.newBuilder() .setScheduleToCloseTimeout(Duration.ofHours(2)) .build()) .build()); public AgentResult delegate(AgentTask task) { return research.research(task); } The endpoint mapping is configured when the caller Workflow implementation is registered, allowing deployment configuration to bind a service name to a Nexus Endpoint without changing business logic. Temporal recommends a collocated pattern by default, where Nexus handlers live beside the Workflows they expose, and a router-queue pattern when routing requires separate scaling, permissions, or deployment ownership. Operationally, Nexus creates a stronger debugging surface than an opaque chain of HTTP requests. Caller histories record Nexus scheduling, start, completion, failure, timeout, and cancellation events. Bidirectional links connect caller events to handler Workflow histories, pending operations expose retry state, and OpenTelemetry integration can visualize call graphs across Nexus Operations, Activities, and Child Workflows. Security follows the service boundary. In Temporal Cloud, Nexus Endpoints use caller-Namespace allowlists, Workers authenticate with mTLS or API keys, and cross-Namespace Nexus traffic is secured by the platform. Nexus payloads use the same Data Converter model as Workflows and Activities, including codec-based encryption. Cloud Nexus Endpoints are not general public HTTP endpoints; they are reached through Temporal SDK execution inside the account. Durable Agent Services, Not Just Durable Requests Temporal Nexus does not replace agent interoperability protocols such as A2A. A2A standardizes how independent agentic applications discover capabilities, delegate tasks, and exchange results, while Nexus connects Temporal applications through durable operations. The boundary is complementary: an external A2A-facing layer can provide ecosystem interoperability, while Temporal Workflows and Nexus provide durable execution for agent services running inside Temporal. The deeper shift is from treating agent delegation as message delivery to treating it as durable work. HTTP can move a request across a network, but production agent collaboration also needs remembered state, bounded retries, deduplication, cancellation propagation, durable completion, access control, and traceable execution across service boundaries. Nexus puts those semantics directly into the invocation model. For long-running or side-effecting agent systems, that changes the handoff from an unreliable gap between services into a first-class execution boundary that can survive process crashes, worker outages, and delayed completion without losing the thread of the operation. More
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery

Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery

By Uthej Mopathi DZone Core CORE
A self-healing SQL pipeline should not mean autonomous SQL generation followed by privileged execution. In production, the safer interpretation is narrower, where a language model proposes a repair, while deterministic controls decide whether that repair is syntactically valid, semantically plausible, operationally safe, and eligible for execution. This distinction matters because the same mechanism that corrects a renamed column can also generate an unintended DELETE, widen a join, or scan an unexpectedly large dataset. Structured-output features can constrain an LLM response to a defined schema, but schema conformance is not equivalent to database correctness or authorization. OpenAI’s Structured Outputs is designed to make generated output conform to supplied JSON Schemas, and it does not validate SQL semantics or execution safety. Treat the Failure as Evidence, Not Merely a Prompt The repair loop should begin by classifying the failure before any model call. Parser errors, missing relations, unknown columns, type mismatches, permission failures, timeouts, cardinality explosions, and upstream freshness problems require different responses. A permission error should not trigger SQL rewriting, while an unknown-column error may justify metadata inspection. This gate keeps deterministic failure classes deterministic. Consider a pipeline that previously executed: SQL SELECT customer_id, customer_segment FROM analytics.customer_profile WHERE active = TRUE; An upstream migration renames customer_segment to segment_name. The database returns an unknown-column error. The repair service should collect the failed statement, SQL dialect, error code, schema version, referenced objects, and recent catalog changes. Metadata inspection can then establish that customer_profile still exists, customer_segment no longer exists, and segment_name appeared in the latest schema version. That evidence is stronger than asking a model to infer a replacement from exception text alone. Schema drift detection should therefore precede generation. Catalog snapshots can be hashed and compared between successful and failed runs. Candidate mappings can incorporate data type compatibility, nullability, lineage metadata, column comments, and migration records. The LLM then receives only evidence relevant to the suspected failure class. Candidate generation should return a structured proposal rather than free-form SQL. A production contract can require the proposed SQL, repair category, changed identifiers, evidence references, and assumptions: Python def generate_candidate(failure, metadata): return llm.generate( schema=RepairProposal, context={"failure": failure, "metadata": metadata}, constraints={"max_statements": 1, "allow_dml": False} ) The allow_dml flag is an application policy, not an instruction trusted merely because it appears in a prompt. Structured generation narrows output shape, while authorization remains outside the model. Anthropic’s evaluation guidance similarly emphasizes measurable success criteria and testable thresholds rather than treating model behavior as inherently reliable. Parse the Candidate Before the Database Sees It String matching is too weak for SQL safety. A rule such as "DELETE" not in sql.upper() can miss nested statements and dialect-specific constructs. The candidate should be parsed into an abstract syntax tree using the correct database dialect. SQLGlot parses SQL into expression trees and supports multiple dialects, enabling structural inspection before execution. A validator can reject statement types outside an allowlist and verify referenced tables and columns against current metadata: Python def validate_ast(sql, dialect, catalog): tree = parse_one(sql, dialect=dialect) if tree.find(Delete) or tree.find(Update) or tree.find(Insert): raise PolicyViolation("mutating statement rejected") for table in tree.find_all(Table): catalog.require_table(table.name) for column in tree.find_all(Column): catalog.require_column(column.table, column.name) return tree AST validation should also enforce tenant boundaries, prohibited schemas, mandatory predicates, join limits, and function restrictions. A generated query can be syntactically valid yet unsafe because an omitted filter changes a targeted lookup into a full table operation. Semantic checks therefore need context from the original successful query, expected output columns, and data quality assertions. The repaired query may be: SQL SELECT customer_id, segment_name AS customer_segment FROM analytics.customer_profile WHERE active = TRUE; Preserving the original output alias matters because downstream consumers may depend on customer_segment even though the physical source column changed. A repair that only substitutes the new identifier could restore execution while silently breaking the pipeline contract. Dry Runs Should Prove More Than Syntax A candidate that survives static validation still should not immediately reach production data. Database-native validation can catch failures that an AST cannot. BigQuery dry runs validate query structure and estimate bytes processed without executing the query, making them useful for rejecting unexpectedly expensive repairs. Google also documents that successful dry runs do not guarantee successful runtime execution and that multi-statement dry runs have special limitations. A guarded validation step can combine dry run results with policy limits: Python def dry_run(candidate): result = warehouse.validate(candidate.sql) if result.bytes_scanned > MAX_BYTES: raise PolicyViolation("scan budget exceeded") if result.output_schema != candidate.expected_schema: raise ContractViolation("output schema changed") return result For engines without native dry run support, a read-only transaction, isolated replica, sandbox database, or planner-only operation can provide a safer boundary. PostgreSQL supports read-only transaction modes that prevent changes to non-temporary tables, adding a database-enforced control rather than relying only on application logic. Operational safety also depends on continuous monitoring after a repair is deployed. Query latency, row counts, null rates, schema changes, and downstream data quality indicators can reveal subtle regressions that static validation may miss. Post execution monitoring therefore provides another deterministic checkpoint, allowing suspicious behavior to trigger rollback or human review immediately. Confidence should come from independently observable signals, not from an LLM declaring confidence in its own answer. A score can combine schema evidence, AST-policy results, dry run success, output-schema stability, and historical repair success: Python def repair_score(signals): return ( 0.30 * signals.schema_evidence + 0.25 * signals.ast_validation + 0.20 * signals.dry_run + 0.15 * signals.contract_match + 0.10 * signals.history ) Weights should be calibrated against labeled historical failures. LLM-based judging can contribute a secondary semantic signal, but it should not authorize execution. G-Eval shows that model-based evaluators can correlate with human judgments while also identifying evaluator bias as a concern. A useful inference from RAGAS is that component-level measurements are more diagnosable than one opaque score; the same principle fits SQL repair by keeping schema, syntax, execution, and contract evidence separately observable. Bound Autonomy and Make Every Repair Reversible Self-healing becomes dangerous when retries are unbounded. A failed candidate should feed only new deterministic evidence into the next attempt, such as a parser error or dry-run diagnostic, and the loop should stop after a small configured limit. Repeated failures, low confidence, ambiguous schema mappings, contract changes, or any proposed mutation should escalate to human review. LangSmith distinguishes offline evaluation from online production evaluation and describes a feedback loop in which production failures become future evaluation cases and the same operational pattern fits SQL repair systems. Every attempt should produce an immutable audit record containing the original SQL hash, failure evidence, metadata version, model and prompt version, candidate hash, validation outcomes, confidence signals, execution identity, and final disposition. That record supports debugging and regression testing of future repair policies. Production monitoring should also track changes in failure distributions and repair success rates, as Google’s MLOps guidance treats monitoring as a trigger for new experimentation when production quality degrades. Automatic mutation deserves a higher bar than automatic read-only repair. When writes are permitted, idempotency keys should prevent duplicate side effects across retries, and execution should remain transactional whenever supported. AWS reliability guidance recommends idempotency for database insert, update, and delete operations. Transaction savepoints and rollback mechanisms add another containment layer, as PostgreSQL savepoints allow effects after a savepoint to be selectively discarded without abandoning the entire transaction. A production-grade self-healing SQL pipeline is not an autonomous database administrator implemented with a prompt. It is a controlled repair system in which probabilistic generation is surrounded by deterministic evidence collection, AST inspection, schema validation, database-enforced dry runs, calibrated confidence thresholds, bounded retries, auditability, and explicit escalation. The safest design grants the LLM authority to propose change, not authority to approve or execute it. With that separation preserved, LLMs can reduce recovery time for routine SQL failures while the database, policy engine, and human review path retain control over production state. More
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
By DZone Staff
Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival
Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival
By Ai Cerrudo

Refcard #291

Code Review Core Practices

By Vidyasagar (Sarath Chandra) Machupalli FBCS DZone Core CORE
Code Review Core Practices

Refcard #267

Getting Started With DevSecOps

By Akanksha Pathak DZone Core CORE
Getting Started With DevSecOps

More Articles

Predict, Repeat, Improve: Deterministic Simulation Testing Explained
Predict, Repeat, Improve: Deterministic Simulation Testing Explained

It’s 2 AM. Your phone buzzes, the on-call alert flashes, and suddenly you are staring at a production outage that makes no sense. Following the logs, you get a hint: when you rerun the same scenario in staging, everything behaves perfectly. None of the quality gates/QA pipelines catch it, chaos experiments didn’t reproduce it, and now a ghost is chased that only appears when the system is under real-world pressure. Distributed systems are notorious for these “phantom failures” — rare timing-dependent bugs that surface unpredictably and vanish just as quickly. They are the kind of dreaded incidents that keep engineers awake at night because they are unreproducible. Take a real-world example: A service once crashed because two nodes tried to become leader at the exact same millisecond. In staging, the timing never aligned that way, so the bug remained invisible. But in production, under heavy load, it just happens, sending the system into chaos. Engineers spent days trying to recreate the failure, but without a deterministic replay, it's pure luck to get a reliable reproduction. Even when it happens, engineers may not be sure what caused it or how to reproduce it deterministically. Enter deterministic simulation testing (DST). DST builds a fully controlled, re-playable simulation of your system’s world — nodes, clients, clocks, network delays, partitions — so that even the most elusive bugs can be identified, captured, replayed, and studied. In this article, we will uncover how DST can transform those unpredictable 2 AM incidents into predictable, debuggable coordinates — giving you a new way to tame the chaos of distributed systems. Deterministic Simulation Testing Definition Deterministic simulation testing (DST) is a software testing methodology that places the system under test within a fully controlled, simulated environment. All sources of non-determinism — system clock, thread scheduling, network, disk — are intercepted and made deterministic. The key property is that, for a given initial seed and configuration, the entire execution is reproducible. The same sequence of events, faults, and outcomes will occur on every run with that seed. Key Concepts Breaking down the above definition, below are the key concepts for DST: Determinism → The system’s behavior is purely dependent upon its initial state and the seed. All non-deterministic sources are simulated to achieve determinism.Simulation → The system is run in a virtual environment that can simulate faults and control the passage of time. E.g., controlled clock skew introduced across various nodes, added network delays to achieve out-of-order event delivery.Reproducibility → Any failure or bug found during simulation can be reliably reproduced by rerunning the simulation with the same seed.Scenario Exploration → By varying the seed and/or simulation parameters, DST systematically explores a vast range of possible execution paths and failure scenarios. How DST Works Let's consider a simple scenario where two users update and read the same record in a very short interval. User 1 updates a record with a new value at time instance T0, and User 2 reads the same record at time instance T1. Note that the interval between T0 & T1 stays the same. In an ideal case (i.e., scenario 1), the new value is updated or written immediately, i.e., without any delay. Thus, when User 2 reads the same record at T1, it is able to read the latest value. In scenario 2 suppose the write is delayed due to network partitioning, disk write etc. User 2 thus sees the old value of the record even if it reads the value at the same time instance T1. Although a stale read may look trivial, it may lead to workflow halt, process crash, etc. in a complex real-world system. Imagine what could happen in a real-world system where: Multiple processes are scheduled for execution, within and across nodes.Multiple network calls are made between several nodes.Multiple operations are performed by several distributed processes on a single disk. Traditional testing strategies or frameworks are inherently constrained and thus can’t simulate such delays or faults. Because of this, it's nearly impossible to identify, catch, reproduce, or debug issues arising from such situations — rendering them unreliable or, at best, non-deterministic. To achieve determinism, the testing framework must take total control over the environment to intercept and manage all external interactions as described below: Controlled scheduling → Instead of relying on the operating system’s unpredictable thread/coroutine scheduler, the simulator provides its own deterministic scheduler. Thus ensuring various scheduling combinations are simulated.I/O mocking → All network calls, disk writes, and clock queries are routed through the simulator, allowing it to inject latency, drop packets, or change the time (e.g., clock skew).Single-threaded execution → Many DST frameworks run the entire distributed system stack within a single thread, completely stripping away the chaotic, unrepeatable nature of multi-threading. Thus, by eliminating real-world “flakiness,” DST allows developers to reproduce chaotic distributed system bugs with perfect precision, thanks to its inherent ability to replay any failing execution: Seed-based replay → The same seed reproduces the exact sequence of events, making debugging tractable.Time-travel debugging → Some platforms (e.g., Flashback) allow stepping backward and forward through execution, inspecting state at any point for a granular view of the system. DST Implementation Approaches and Architecture Patterns Below are two approaches for DST. Pluggable Non-Determinism Design the system so that all non-deterministic components (clocks, I/O, etc.) are pluggable. This strategy is used by TigerBeetle. This requires: Abstracting all system interactions behind interfaces.Providing both real and simulated implementations.Ensuring that the same codebase can run in both production and simulation by swapping implementations at startup. Pros: Deep control and minimal divergence between test and production code. Suitable for greenfield systems. Cons: Requires significant upfront design and is challenging to retrofit into existing systems. Deterministic Hypervisors and Emulation A more recent and flexible approach is to run unmodified binaries inside a deterministic hypervisor or emulation layer. This strategy is used by Hermit and Weave. Pros: Can test existing systems without code changes; language-agnostic; simulates the entire stack. Cons: May have performance overhead; some system behaviors may escape determinism if not fully intercepted. Benefits System employing DST benefits as below: Identify and reproduce rare failures → DST allows engineers to replay the exact sequence of events that led to a bug. This eliminates the frustration of “flaky” issues that appear inconsistently, making debugging far more reliable. Moreover, DST helps find bugs in execution paths unreachable by example-based tests.Improved developer productivity → Bugs are easier to reproduce, debug, and fix; less time spent on war rooms and emergency triage.Improved confidence in correctness → DST validates critical invariants (like consensus, failover, or transaction consistency) under controlled simulations. Engineers gain assurance that core distributed protocols behave as expected even under stress. Thus, preventing rare bugs from reaching production, increasing system uptime and user trust.Scalable debugging for complex systems → In microservice or event-driven architectures, DST helps tame the exponential growth of possible interleavings by focusing on deterministic seeds. This makes large-scale debugging more tractable. Challenges and Limitations While DST provides unparalleled confidence, it requires significant architectural investment. Retrofitting DST into existing systems may require significant refactoring. It can be highly intrusive, requiring developers to write custom code or frameworks, as production code often cannot rely on external third-party libraries that invoke un-mocked I/O or system calls. Moreover, DST requires careful modeling of external systems to avoid missing integration bugs. Ensuring sufficient coverage without combinatorial explosion is a major challenge — especially in modern systems with multiple integration points. Below are gaps in tooling that limit DST outcomes: Language and platform support → Not all languages and runtimes have mature DST frameworks.Hypervisor limitations → Deterministic hypervisors may not support all system calls or hardware features. DST Comparison and Applicability DST vs. Chaos Engineering DST is proactive and enables perfect reproducibility. It is best suited for development and pre-production, catching bugs before they reach users. It can simulate production chaos in minutes, and every failure is a permanent regression. Chaos engineering is reactive, non-deterministic, and validates the behavior of real deployments. It is essential for catching issues arising from real infrastructure, misconfigurations, or dependencies that simulation cannot model. However, it cannot guarantee coverage or reproducibility, and carries the risk of impacting users. In essence, both DST and chaos engineering are complementary to each other and are necessary for comprehensive reliability. DST in Functional vs. Performance Testing Functional Testing DST is ideally suited for functional testing. Validates correctness under all possible interleavings, failures, and workloads.Checks invariants, safety properties, and liveness under stress.Finds rare, timing-dependent bugs that are invisible to example-based tests. Performance Testing DST is not primarily designed for performance testing. The simulated environment does not reflect real hardware performance, network latency, or throughput.Time is virtualized and compressed; I/O is in-memory.Performance metrics (latency, throughput) measured in simulation may not correspond to real-world values. However, DST can be used to: Validate performance-related invariants (e.g., absence of deadlocks, progress under load).Simulate pathological scenarios (e.g., extreme contention, resource exhaustion) to observe system behavior. Recommendation Combine DST for functional correctness with real-world performance and benchmarking suites for comprehensive validation. DST Applicability to AI/ML and Agentic Systems AI/ML systems, especially those based on large language models (LLMs) and agentic workflows, are fundamentally non-deterministic. This makes traditional testing and debugging extremely difficult, with “heisenbugs” that vanish when observed. DST can be adapted to AI/ML systems by creating controlled, simulated environments for agents to operate in. Or using a hybrid approach of combining deterministic components (rule-based logic) with LLM-driven reasoning, using seeds to replay failures. Case Studies FoundationDB, with its deterministic simulator tool, achieved legendary reliability by running trillions of simulated CPU-hours, finding and fixing every known bug before production.TigerBeetle built a Viewstamped Operation Replication simulator (VOPR) to simulate financial transaction systems, catching subtle bugs in consensus and replication.Ethereum Merge used Antithesis to test the transition to Proof-of-Stake, simulating multiple client implementations in a deterministic environment. Conclusion Deterministic simulation testing (DST) represents a paradigm shift in the testing and validation of distributed systems. By enabling exhaustive, reproducible exploration of the vast state space of concurrent, failure-prone systems, DST empowers engineers to find and fix the rarest and most pernicious bugs before they reach production. Its integration with property-based testing, fuzzing, and fault injection, combined with advances in deterministic hypervisors and simulation frameworks, has made DST accessible to a growing range of systems and organizations. While DST requires significant engineering investment, careful system design, and ongoing maintenance, its benefits in reliability, developer productivity, and user trust are profound. As distributed systems continue to grow in complexity and AI/ML systems become more agentic and autonomous, the need for rigorous, deterministic validation will only intensify. The future of DST lies in deeper integration with formal methods, smarter state-space exploration, and broader applicability to AI/ML and hybrid systems. Organizations that embrace DST, alongside complementary techniques like chaos engineering and formal verification, will be best positioned to deliver robust, trustworthy, and resilient distributed systems in the years ahead. DST Tools and Frameworks Deterministic simulation testing (DST) tooling is still a niche but growing ecosystem. Each has a unique focus — ranging from language-level deterministic runtimes to full-stack hypervisor-based reproducibility. Based on the specific needs a single or combination of them can be picked up. References and Further Reads Taming Chaos — DSTSquashing the Heisenbug with DSTAntithesis — DSTPhil Eaton — DSTRedstone — DST FrameworkResonate — DSTJespen | TickLoom

By Ammar Husain DZone Core CORE
Detection and Response Did Its Job. Now Someone Has to Actually Fix It.
Detection and Response Did Its Job. Now Someone Has to Actually Fix It.

A host is isolated. A malicious process is killed. A compromised credential is revoked before an attacker can use it again. Detection and response worked exactly as intended, automatically, correctly, and fast. What follows is usually less tidy. The attacker may be gone, but the opening they used can still be sitting there waiting for the next attempt. Veracode’s State of Software Security Report gives some sense of the problem’s scale. Critical security debt was up 20% year over year, and high-risk vulnerabilities, those considered both severe and highly exploitable, climbed 36%. Detection has made progress, but finding a problem quickly and implementing fixes are clearly not the same thing. A cybersecurity platform that’s actually good at threat detection and incident response has to do more than contain the immediate event. The incident still has to be reconstructed, connected to the team responsible for the affected service, and followed back to whatever condition made the attack possible. Only then can developers or IT managers make the change and return to the environment to see whether it worked. Reconstructing What Actually Happened An isolated host does not come with a narrative attached. Neither does a killed process. What the security team initially has is an event, and before someone can fix the underlying problem, they need to reconstruct the sequence that produced it. Sysdig offers a useful example of what that evidence can look like. Its Falco-based real-time detection can sit alongside response actions such as killing a process, pausing a container, or quarantining a file. But stopping the activity is only part of the process. The surrounding runtime evidence can show an investigator what was happening on the system when the alert fired, rather than leaving the team to piece together the incident after the fact. That evidence can include processes being executed, network connections, changes to files, and the lineage between processes. On Linux hosts, Kubernetes nodes, and VMs, Sysdig can collect system-call data using eBPF-based drivers, with kernel modules also supported. Serverless environments require instrumentation appropriate to that execution model, while Windows uses Event Tracing for Windows rather than Linux eBPF for kernel-level workload visibility. What all of this data has in common is that it captures actual behavior, not just another alert. It provides evidence of what software did in production, evidence that becomes the raw material for figuring out what really happened during an incident. That evidence is what runtime security is actually for. Instead of merely producing another alert, it’s most useful when scans reveal what software did in production once something has already gone wrong. Fixing what made that possible is a separate job, which is exactly why runtime and build-time security have to work together rather than standing in for each other. Finding Whose Service This Actually Is Reconstructing the attack is one problem. Figuring out whose problem it is can be another. A workload running in production carries plenty of technical information, but it does not necessarily tell the responder which repository produced it, which engineering team is responsible for it, or who is in a position to change it without breaking something else. That ownership gap becomes particularly difficult in cloud-native environments. A compromised container may belong to a service assembled from several repositories, deployed through shared infrastructure code, and operated by a platform team that did not write the vulnerable application. A credential may technically belong to one cloud account while being consumed by workloads owned elsewhere. Most organizations already have clues scattered across their environment. A service catalog may name an owner, CODEOWNERS may point to a team, repository metadata may identify maintainers, and infrastructure tags or deployment history can help connect what is running back to where it came from. None of those records is particularly useful, though, if it describes an organization that existed six months ago rather than the one responding to the incident today. Ultimately, the security team needs to get from “this workload was compromised, and our automation has contained it” to the much more useful, “we know which team can remediate the issue that made this situation possible.” Tracing Why It Was Actually Exploitable At this stage, the responder knows the shape of the incident and has a reasonable idea of who should own the fix. What is still missing is the cause. The question shifts from what the attacker did to what was present in the environment that let those actions succeed. The process that was killed may only be the last link in a much longer chain. Perhaps an old dependency made it into the container image. Maybe a role had permissions it never needed, a service was reachable from somewhere it should not have been, or an infrastructure configuration quietly exposed a path into the workload. Unless that earlier condition changes, stopping the process deals with the incident without really dealing with its cause. This is the part of the workflow Wiz’s Green Agent is designed to address, the stage most detection and response tooling stops short of. Once a runtime finding has been detected and contained, Green Agent picks up from there, analyzing the confirmed finding in the context of the environment, tracing the issue back to its root cause, identifying ownership, and generating environment-specific remediation guidance for the developer or owner positioned to make the change. The Security Graph provides supporting context across areas such as identities, network exposure, cloud resources, code, and data sensitivity, helping distinguish the immediate runtime event from the condition that made it exploitable, so detection and response extend past containment into an actual fix. Tracing a runtime event back to the code that caused it isn’t a new problem. Application security teams have been working on closing the gap between what SAST catches in code and what DAST catches at runtime for years. A finding is only actionable once it’s connected back to its source. Incident remediation is running into the same wall now, just with an attacker involved instead of a scanner. Getting the Fix Into the Developer’s Workflow Root-cause analysis can still fail operationally if its result lives inside a security console that the responsible developer rarely opens. The handoff therefore matters almost as much as the diagnosis. Security teams need to decide which findings require human intervention and then put those findings into systems where engineering work is already managed. Palo Alto Cortex XSIAM handles the triage side of that decision. Embedded automation enriches alerts and closes low-risk cases before they ever reach an analyst’s queue, leaving higher-value cases for actual investigation and response. The developer who has to make the change probably isn’t working out of a security console at all. Their day is more likely to revolve around a repository, an issue tracker, a pull request, an IDE, and team messages. That reality colors what a useful handoff looks like. Sending another alert is not enough. The developer needs enough of the incident’s story to understand why the change is being requested and enough technical context to know where to start. The lesson is familiar from application testing. Making DAST findings actionable requires narrowing the distance between discovering a vulnerability and getting it into a form developers can realistically act on. Runtime incidents create much the same handoff problem, only with an attacker potentially having demonstrated the consequence already. Confirming the Fix Actually Held The last stage is easy to treat as administrative cleanup, but it is part of remediation itself. A developer changes a dependency, tightens a role, modifies a configuration, or removes an exposure. The original condition then needs to be tested again. A clean retest is useful evidence that the remediation changed the behavior security observed. It is not absolute proof. An endpoint may have moved, an access path may have changed, or another control may now be masking the original condition without eliminating its cause. The same caution applies to incident response. Reconnecting a host or restoring a quarantined workload does not demonstrate that the weakness behind the incident is gone. Without retesting, the organization has confirmed that containment can be reversed, not that remediation succeeded. Conclusion Automated detection and response is extremely effective at handling emergencies. It can interrupt malicious activity faster than a human analyst could reasonably investigate and act. What Veracode’s numbers suggest, however, is that stopping the immediate event has not solved the industry’s bigger challenges. The work that follows is slower and crosses more boundaries, from forensic evidence to service ownership, engineering changes, and finally another look at the environment to make sure the original weakness is no longer there. That full path is what separates a genuinely complete approach to detection and response from an approach that only detects and contains. A response system can stop the attacker and give the organization breathing room, but somebody still has to use what the incident revealed to remove the condition behind it. Then the organization has to check its work. Isolation buys that opportunity, and remediation is what makes use of it.

By Philip Piletic DZone Core CORE
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data

Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq

By Jubin Soni, FBCS DZone Core CORE
How to Test POST API Requests With Playwright TypeScript
How to Test POST API Requests With Playwright TypeScript

Testing POST API requests is an important skill for modern QA and automation engineers working with backend services and microservices. In this article, we’ll explore how to test POST API requests with Playwright and TypeScript, focusing on sending a request body using different approaches. By the end, you will learn how to send POST API requests using the following approaches for adding a request body: JSON Object/Array(data)Stringified JSONJSON FileFaker library Application Under Test We will use the POST /addOrder API of the RESTful e-commerce demo application for this demo. The API schema is provided below: JSON { "user_id": "string", "product_id": "string", "product_name": "string", "product_amount": 0, "qty": 0, "tax_amt": 0, "total_amt": 0 } ] Testing POST API Requests With Playwright TypeScript Playwright provides a powerful request API that allows us to create and manage HTTP request contexts. Let’s walk through how to send POST API requests step-by-step using different approaches for passing the request body: Request Body as JSON Object/Array Let’s send a POST API request with a static JSON Array and verify that the response status code is 201. TypeScript test("POST order details API with static JSON Array", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data:[{ user_id: "1", product_id: "82", product_name: "Cadbury Bar", product_amount: 12, qty: 2, tax_amt: 1, total_amt: 25 }, { user_id: "2", product_id: "80", product_name: "MilkyBar", product_amount: 10, qty: 1, tax_amt: 1, total_amt: 11 } ], }); expect(response.status()).toBe(201); }); Code Walkthrough A Playwright test case for testing the POST API request is created using the request API, which is a built-in Playwright APIRequestContext fixture used to send HTTP requests. Sending the POST Request: The request.post() sends a POST request to the “http://localhost:3004/addOrder/” endpoint. The response received after sending a POST request is stored in the response variable.Passing the Request Body (Static JSON Array): The data property is used to send the request body, which accepts data in JSON Array and JSON object formats. Since the POST /addOrder API accepts an array of order details, allowing multiple orders to be submitted within a single JSON array, the test provides the data in JSON array format.Validating the Response: The response.status() retrieves the HTTP status code. The assertion ensures that the status code 201 is returned, indicating that the orders were successfully created. Request Body Using JSON.Stringify() JSON.stringify() can be used while sending raw payloads, testing malformed JSON, or sending custom-formatted JSON. However, the Content-Type header should be set to application/json so the server correctly interprets the request body as JSON data. Let’s write the same POST /addOrder test using JSON.stringify() and verify that a status code 201 is returned in the response. TypeScript test("POST order details API using JSON.Stringify", async ({ request }) => { const orderData = [{ user_id: "5", product_id: "64", product_name: "Cadbury Mini", product_amount: 5, qty: 3, tax_amt: 1, total_amt: 16 }]; const response = await request.post("http://localhost:3004/addOrder/", { data: JSON.stringify(orderData), headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); The orderData array is defined to contain one order object. Each field represents the order details such as user_id, product_id, product_name, and so on. This is a normal JavaScript object at this stage, not JSON yet. The JSON.stringify(orderData )converts the JavaScript object into a JSON string. As we manually stringify the request payload, we should explicitly set the Content-Type header to application/json. It tells the server to treat the request body as JSON data. Playwright sends the POST API request using the post() method and then validates that the response returns a 201 status code. Request Body as a JSON File Using a JSON file as a request body is a handy approach when testing POST API requests with Playwright TypeScript. It supports large payloads and allows the same file to be easily reused across multiple tests. The following orders.json file will be used as a request payload in the POST /addOrder API: JSON [ { "user_id": "1", "product_id": "79", "product_name": "5 star 10gm Chocobar", "product_amount": 5, "qty": 1, "tax_amt": 0.5, "total_amt": 5.5 }, { "user_id": "2", "product_id": "71", "product_name": "Lindt Milk Chocolate", "product_amount": 15, "qty": 3, "tax_amt": 2.5, "total_amt": 47.5 } ] The following configurations should be in place before we proceed to write the test using a JSON file as the request payload. tsconfig.json JSON { "compilerOptions": { "module": "NodeNext", "moduleResolution": "NodeNext", "resolveJsonModule": true } } We need to ensure that the JSON file is imported first and then attached to the data parameter while sending the POST request. TypeScript import orders from '../test_data/orders.json' with {type: 'json'}; test("POST order details API using JSON file", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data: orders, }); expect(response.status()).toBe(201); }); This test imports test data from an external orders.json file and uses it as the request body for a POST API call in Playwright. The request.post() method sends the imported JSON directly as the payload, and the test verifies that the API returns a 201 status code. Using a JSON file makes it easier to manage, reuse, and update request payloads without modifying the test logic, improving the maintainability and readability of API tests. Request Body With Faker Library Using the Faker library to create a request body for a POST API helps generate realistic and dynamic test data, reduces hardcoded values, and improves test coverage. It is especially useful for simulating real-world scenarios and avoiding duplicate data issues. However, it also has limitations. Since Faker generates random data, it may lead to inconsistent test results if not properly handled, and debugging failures could become more difficult without fixed or reproducible inputs. To use the Faker library, we need to install it first using the following command: Plain Text npm install --save-dev @faker-js/faker Next, let’s create two helper functions: one to create order objects with the required order details, and another to generate an array of orders based on the count provided by the user. TypeScript import { faker } from '@faker-js/faker'; export function createOrderDetails() { const productAmount:number = faker.number.int({ min: 1, max: 100 }); const qty:number = faker.number.int({ min: 1, max: 5 }); const taxAmt:number = faker.number.int({ min: 2, max: 10 }); const totalAmt:number = (productAmount*qty)+taxAmt; return { user_id: faker.number.int({ min: 1, max: 50 }), product_id: faker.number.int({ min: 1, max: 100 }), product_name: faker.commerce.productName(), product_amount: productAmount, qty:qty, tax_amt: taxAmt, total_amt: totalAmt }; } The createOrderDetails() function returns the order details with the required fields. It also calculates the total amount by summing the total product value and tax amount. The values in all fields are updated randomly using the appropriate methods provided by the Faker library. TypeScript import { faker } from '@faker-js/faker'; export function createRandomOrders(count:number) { return faker.helpers.multiple(createOrderDetails, {count}); } The createRandomOrders(count) method accepts a count parameter and returns randomly generated orders that can be directly used in the test. TypeScript return faker.helpers.multiple(createOrderDetails, {count}); This is where the work happens. multiple() is a helper function provided by Faker. Its job is to execute another function multiple times and collect all the results into an array. It takes two arguments: A function to execute repeatedly.An options object specifying how many times to execute it. It asks Faker to call createOrderDetails() exactly count times, collect all the generated orders into an array, and return that array to the caller. TypeScript test('POST order details API using Faker library', async({request}) => { const orderData = createRandomOrders(5); const response = await request.post("http://localhost:3004/addOrder/", { data: orderData, headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); This test sends a POST request to the /addOrder endpoint API using dynamically generated test data from the Faker library. The createRandomOrders(5) function generates an array of 5 random order objects, which are passed as the request body using the data property. The test then verifies that the API responds with a 201 status code, confirming that the orders were successfully created. Summary In this tutorial, we explored multiple approaches to testing POST API requests using Playwright with TypeScript, including sending static JSON payloads, using JSON.stringify, importing data from external JSON files, and generating dynamic test data with the Faker library. We also covered how to handle headers correctly and validate API responses using status code assertions. In my experience, using external JSON files is a practical approach for testing POST API requests because it lets us send bulk data from a single file. Similarly, the Faker library can also be used to generate dynamic test data; however, in some cases, using a third-party library may not be permitted. Ultimately, the software team chooses the approach that best fits their testing strategy. Happy testing!

By Faisal Khatri DZone Core CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1

A few months ago, one of our teams celebrated a milestone in their quarterly review. AI adoption was up. The productivity dashboard showed that developers were generating, on average, 40% more code per sprint. The tech lead showed this as a major win. Three weeks later, I was on a call for a production issue. The challenge was that errors were coming in three different formats depending on which endpoint you hit. The alerting was blind to a category of failure it had always caught before. We had a centralized exception handler. It logged context, mapped it to the right HTTP status, and pushed these alerts to our observability stack. When we investigated, we found that recent AI-assisted PRs had started introducing their own try-catch blocks inline. Each one caught exceptions locally, logged in a slightly different format, and returned a slightly different error shape. Some swallowed the exception instead of letting it propagate up to the handler that would have alerted us. Each one of those PRs was correct. Every one passed review, including the reviews I did myself. We were all checking for correctness, and consistency isn't the kind of thing that shows up in a git diff. Cleaning it up took most of a sprint. And that productivity dashboard? It counted the original generation and the cleanup as output. Twice the code, twice the "productivity," for a net loss of engineering time. The dashboard was going up while the system got worse underneath it. That's the mistake I keep seeing where teams confuse code generation with engineering progress. The LOC Trap, Reloaded Fred Brooks called this out in The Mythical Man-Month decades ago. He explicitly mentioned that measuring programming productivity by lines of code is nonsensical. Everyone agreed and then somehow forgot. So why did we rebuild the exact same dashboard the moment AI arrived? We are using the same flawed metric, with AI branding, and presenting it to boards. When a team lead reports that AI tools helped produce 40% more code, the follow-up I want to hear is “Did we actually need 40% more code?” Usually the answer is no. What we needed was the same outcomes with less effort, and effort in software lives overwhelmingly outside the act of typing. I’ll come back to that. There is a difference this time, and it is worth naming. Teams aren’t defending line counts out loud anymore. The dashboards have moved on to merged pull requests, agent tasks completed, and suggestions accepted. It is the same instinct in a unit that sounds more respectable in front of a board, counting the artifacts of work and reporting the count as progress. Why Senior Engineers Delete Code Here’s a pattern you’ll recognize if you have led engineering teams for any length of time. Your best engineers often produce fewer lines of code than anyone else on the team. Some of their most impactful weeks come out to negative line counts. That’s expertise rather than laziness. I haven’t fully worked out why the instinct for deletion over addition takes years to develop, but it does. GitClear has been measuring this rather than speculating about it. Their 2026 analysis covers 623 million changes from 2023 to 2026, and the figure that stopped me had nothing to do with volume. Moved code, their proxy for refactoring, dropped from 21% of all changes in 2022 to 3.8% by the middle of this year. Duplicated blocks are up 81% across the same window. Cross-file function calls, which is roughly what reuse looks like in a diff, are down 35%. Those three together describe a codebase that has stopped being rearranged when work goes in, and very little gets moved or deleted. When someone takes 2,000 lines of tangled logic and replaces it with 200 clean ones, that looks like a loss on any volume metric. It is an enormous win for the system, and it is exactly the activity that has gone quiet. One correction I owe, since I quoted the earlier version of this research at people for the better part of a year. GitClear’s 2024 report predicted two-week code churn would double in the AI era. It didn’t double. It went up 15%. The headline projection was too aggressive, and the part almost nobody quoted — refactoring falling off a cliff — turned out worse than predicted. Senior engineers get this intuitively. Every line of code is a liability, because every line has to be read, understood, tested, and maintained. So the best solution often makes code disappear, and that takes different forms like a well-chosen abstraction that kills duplication or a config change that removes a custom implementation. Sometimes it’s just a conversation with the PO that drops the requirement entirely. Now think about what AI coding metrics would say about this. An engineer spends a day understanding a system, realizes three services can collapse into one, and deletes 4,000 lines. By every AI productivity metric in use today, that engineer had a terrible day. In reality, they may have saved the organization months of future pain. Gergely Orosz tells a revealing story about what happens when you optimize for the wrong signal. When Uber introduced diff-count metrics, engineers started creating more, smaller changes to look productive. They flooded CI systems, driving up costs. The metric improved and engineering got worse. We are setting ourselves up for the same trap with AI-generated LOC. AI’s Tendency Toward Verbose Implementations This gets worse when you look at what AI coding tools actually excel at, which is producing plausible code quickly. That skill carries a built-in bias toward verbosity. Sometimes more code is genuinely the right call. Explicit beats implicit. A verbose but readable implementation can be better than a clever one-liner that nobody understands at 3 am when production is on fire. I’m not arguing for code golf. But AI-generated verbosity is a specific kind of bad, because it is default verbosity from ignorance of context rather than chosen verbosity for clarity. This distinction matters a lot more than I initially thought. Ask an AI assistant to implement a feature, and you’ll get a complete, working solution, longer than what an experienced developer would write, because it optimizes for correctness and completeness in isolation. It may not pick up the utility you wrote last month to do exactly this. It doesn’t realize the framework provides a one-liner if you structure the problem slightly differently. It can’t tell the difference between “I should be explicit here for readability” and “I am reinventing something that already exists three directories over.” The handler drift I opened with is the cleanest example I have of it. The AI did exactly what it was asked, every single time. Each PR added code, each one passed review on its own terms, and nothing was wrong inside any of them. What we lost lived across them. One handler gave us one error shape, and one error shape gave our alerting something to fire on. No single diff broke that, and all of them together did. GitClear tracks error-masking constructs, which is the failure mode buried in there, and they are up 47% since 2023. Our inline handlers were exactly that. I’d like to think we were an unlucky outlier, and the data says we were ordinary. I’m still not sure how you review for a property that isn’t visible in the file in front of you. Complexity as the Hidden Cost Code volume isn’t a perfect proxy for system complexity, and I acknowledged that above. But it is a directional one, and in aggregate it holds. When your codebase grows by 30-40% in a quarter without a corresponding growth in functionality, complexity is almost certainly growing with it. And system complexity is the single biggest thing determining how fast your team can move over time. I haven’t found a way around that in eighteen years. Every line of code carries ongoing costs that nobody puts on a dashboard: Cognitive load for anyone working nearbyTest coverage, without which it becomes a ticking riskReview time on every future change to itDependencies it drags in that need constant updatingMigration effort during every platform change When AI tools grow your code volume by 30-40%, each of these costs grows. The productivity gain at the moment of writing is real, and I’m not denying that. But it can be entirely eaten up by the downstream cost of maintaining a bigger, more complex system. Then there is the study I keep coming back to, and the update to it that I nearly missed. In mid-2025, METR found that experienced developers using AI coding tools took 19% longer on real-world tasks. The setup was specific. It had 16 open-source contributors working in their own repositories, code they’d lived in for years, 246 real issues averaging about two hours each, with Cursor Pro and Claude. In February 2026, METR published an update that takes a fair amount of it back. They are redesigning the experiment, and the reasons aren’t flattering to the original. Developers who would no longer work without AI declined to take part at all. Somewhere between 30% and 50% of participants avoided submitting exactly the tasks where they expected to want AI. The pay rate for the follow-up work dropped from $150 an hour to $50, which made recruitment worse again. Their own summary is that the newer data amounts to very weak evidence in either direction, and that developers are probably more sped up now, in early 2026, than the early-2025 estimate suggested. So the 19% was never a fact about AI-assisted development. It was a measurement of sixteen people in one setting, and the people who ran it now think it read low. What survives is the part that makes me think harder. Those developers believed they were 20% faster. Whatever the true effect was, it wasn’t the effect they perceived, and not one of them could feel the gap while it was happening. Reviewing and integrating suggestions consumed time none of them accounted for. Selection bias moved the headline number around, but it does not explain away a room full of experienced engineers being wrong about their own week. I could never square that finding with my own experience, because I do feel faster on certain tasks. The update moves me off the fence, slightly. Maybe I wasn’t fooling myself. I would still put no weight at all on my own estimate of how much faster I am, and that’s close to the only thing here I am confident about. None of that is an argument against the tools, only against measuring them by how much code they produce. The Strongest Number Against Me If you want to argue the other side, the best evidence available today is Microsoft’s. Early in 2026, they rolled Claude Code and GitHub Copilot CLI out across the organization and studied what happened. Tens of thousands of engineers, four months, and the ones who adopted merged roughly 24% more pull requests than the counterfactual said they would have. That is not a lab, and it is not a small group of volunteers. It’s the largest measurement of agentic coding tools anyone has published, the effect is large, and it points the right way. I take it seriously. I also notice what the unit is. The authors get there ahead of any critic. Their paper says a merged PR isn’t the same as the value it delivers, which is the argument of this entire post, conceded inside the study that is supposed to answer it. Twenty-four percent more merged pull requests is consistent with 24% more delivered value. It is equally consistent with the same work arriving in smaller slices, which is what happened at Uber the moment diff count landed on a dashboard. There are narrower caveats, and I won’t pretend I have chased all of them. The comparison is against engineers who already had AI in their IDE, so what it measures is the increment from adding an agent rather than the effect of AI from zero. Four months is not long enough for maintenance cost to turn up. And engineers chose for themselves whether to adopt. None of that makes the study wrong. It is a good measurement of pull request volume, and pull request volume behaves the way lines of code always did. It is easy to report and showcase, if somebody decides that moving it matters. Blind to what the system underneath is doing. What This Actually Means If You Are Leading a Team If you’re six months into your AI investment and your main evidence of ROI is that the output counter went up, whether that counter says lines or pull requests or tasks completed, that should worry you more than it reassures you. You might be measuring the accumulation of future cost and calling it present-day value. Here is what I would look at instead. Cycle time as value reaching production sooner, not just code getting written faster? Those two come apart more often than anyone expects. Rework rate counts if you are fixing more bugs in AI-assisted code? If generated code carries a higher defect rate, your productivity gain is a mirage. Cognitive complexity trend is when your application is getting harder to understand. Tools like SonarQube measure this. If complexity is climbing faster than it did before AI adoption, you’ve got a compounding problem. Developer effort distribution is where the time is actually going. If writing dropped from 20% to 10% of developer time while code review grew from 15% to 30%, you have moved the burden rather than reduced it. None of that is exotic, and it is not just my read. DORA’s 2026 work on the ROI of AI-assisted development arrives at something similar from a different direction. The return on these tools tracks the strength of the engineering system around them rather than the tools themselves. The things that decide it are unglamorous. Whether code review has any slack left in it. Whether people trust the test suite enough to act on a red build, which is the one I have never seen anybody audit. How much sits between a merge and production. Where those are weak, faster generation fills a queue and waits there, and the measurement science is still catching up to the tooling. The Uncomfortable Question Here is what I would ask any engineering leader who reports AI productivity gains based on code output: “If your best engineer spent last week deleting 3,000 lines of AI-generated code and replacing them with 300 lines that do the same thing better, would your dashboard show that as a win or a loss?” On a line count, that week is a catastrophe. On a PR count, it is one merged pull request, which is what a typo fix is worth too. Neither number has any way of seeing what actually happened. If your dashboard shows a loss, you are measuring the wrong thing, and you are rewarding your team for building a larger, slower, more fragile system in exchange for a chart that goes up and to the right. The goal was never more code. It was better systems that deliver business value and are maintained cheaply. Next in this series: Optimizing Benchmark Tasks Instead of Real Delivery Work — why the fact that coding is not the bottleneck makes most AI productivity claims irrelevant to actual delivery speed.

By Gaurav Gaur DZone Core CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads

Batch processing remains vital because many business operations aren't suited to interactive requests. Tasks such as recalculating prices, reconciling transactions, migrating records, generating reports, processing invoices, reclassifying customers, or applying rules across millions of records may require considerable time. Handling these as standard requests leads to fragile systems, increased user wait times, frequent timeouts, challenging retries, and possible data inconsistencies. A batch model handles large workloads predictably, incrementally, and with control over progress and recovery. Rather than processing a massive operation as a single loop, batch processing uses jobs, steps, chunks, checkpoints, filtering, and restartability. This approach separates long-running data tasks from the user experience while delivering a structured execution model. In this article, we will focus on Jakarta Batch and examine its sustained relevance for modern enterprise applications. Why Batch Processing Still Matters in Enterprise Systems Modern applications offer various methods for background processing, such as message queues, event-driven architectures, schedulers, reactive pipelines, and distributed stream-processing platforms. While each addresses specific needs, batch processing is most effective when operations have a defined start and end, involve a known or discoverable dataset, and require controlled execution, progress tracking, restartability, or periodic processing. Batch processing remains essential in enterprise systems. Workloads such as financial reconciliation, billing, payroll, reporting, data migration, regulatory processing, catalog updates, and large-scale reclassification are still prevalent. In these scenarios, the priority is to process large volumes of work safely and predictably, rather than responding to individual events quickly. Batch provides a model specifically designed for these requirements. How Jakarta Batch Works Jakarta Batch organizes background processing into jobs and steps. A job defines the overall batch operation, while each step represents a specific stage. In chunk-oriented processing, a step follows a simple pipeline: read, process, write, and repeat until it processes all input. The Jakarta Batch runtime manages this lifecycle so application code can focus on reading, transforming, and persisting data. A job is the top-level unit of execution and represents a complete business operation, such as importing records, recalculating customer classifications, processing invoices, or reconciling transactions. Jobs can accept parameters at startup, allowing the same batch definition to run with different inputs or business rules. A job consists of one or more steps, each representing a distinct phase of the workload. Simple jobs may have a single step, while complex processes can use multiple steps in sequence, such as importing data, validating it, and generating a final report. Within a chunk-oriented step, the ItemReader supplies data to the runtime one item at a time, from sources such as a database or file. The reader only retrieves the next item and does not need to know how it will be processed or persisted.The ItemProcessor receives each item and applies business rules, such as validation, transformation, classification, enrichment, or filtering. It may return a modified item or null if the item should be excluded from writing.The ItemWriter receives processed items and persists or exports them. Unlike the reader and processor, which handle items individually, the writer typically receives a group of items from the current chunk. This enables more efficient database or bulk operations. Jakarta Batch adds features around this pipeline to support enterprise workloads. The runtime manages chunk boundaries, transactions, checkpoints, execution status, failures, and restart behavior. Chunk size determines how much work is grouped before a write and checkpoint, making it a key parameter for balancing throughput, memory usage, database cost, and recovery. The core model is straightforward: Job → Step → Read → Process → Write → Repeat Jakarta Batch keeps the business pipeline simple while the runtime manages the execution mechanics needed for reliable, long-running data processing. The Sample: Customer Segmentation with Jakarta Batch This example demonstrates the Jakarta Batch model using an e-commerce customer segmentation scenario. Customers are assigned to tiers such as Bronze, Silver, Gold, and Platinum based on configurable spending thresholds. When thresholds change, the application reevaluates the customer base and updates only customers whose classification has changed. The full application includes MongoDB integration, a Jakarta Faces UI, a preview workflow, validation, and supporting services. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. This section focuses on the classes directly involved in Jakarta Batch execution. Starting the Batch Job The application initiates the batch process through CustomerSegmentationService. Unlike the reader, processor, and writer, this class is not a batch artifact. Instead, it is an application service that retrieves Jakarta Batch’s JobOperator from BatchRuntime to start and monitor job executions. Java @ApplicationScoped public class CustomerSegmentationService { public static final String JOB_NAME = "customer-segmentation"; private volatile CustomerSegmentationPolicy currentPolicy; // initialization and status methods omitted public long start(CustomerSegmentationPolicy policy) { if (isRunning()) { throw new IllegalStateException( "A customer segmentation batch is already running"); } Properties parameters = new Properties(); parameters.setProperty( CustomerSegmentationPolicy.JOB_PARAMETER, policy.toJson()); long executionId = BatchRuntime.getJobOperator() .start(JOB_NAME, parameters); currentPolicy = policy; return executionId; } public boolean isRunning() { // implementation omitted } } The key API here is JobOperator, which Jakarta Batch provides as the interface for starting, stopping, restarting, and inspecting jobs. In this example, the segmentation policy is serialized into the job parameters to ensure each execution gets the correct business rules. Reading the Input The first batch artifact, CustomerItemReader, extends Jakarta Batch’s AbstractItemReader to implement a chunk-oriented reader. Java @Named("customerItemReader") @Dependent public class CustomerItemReader extends AbstractItemReader { @Inject private CustomerRepository customerRepository; private List<Customer> customers = List.of(); private int nextIndex; @Override public void open(Serializable checkpoint) { try (Stream<Customer> customerStream = customerRepository.findAll()) { customers = customerStream .sorted(Comparator.comparing(Customer::getId)) .toList(); } nextIndex = checkpoint instanceof Integer index ? index : 0; } @Override public Customer readItem() { if (nextIndex >= customers.size()) { return null; } return customers.get(nextIndex++); } @Override public Serializable checkpointInfo() { return nextIndex; } } These methods are part of the Jakarta Batch reader lifecycle defined by AbstractItemReader. open() prepares the reader and accepts a previous checkpoint if available. readItem() provides the next item to the runtime; returning null indicates there is no more input. checkpointInfo() reports the reader’s current position for checkpointing. For simplicity, this sample loads customers into memory. For larger workloads, the implementation might use pagination or a MongoDB cursor without changing the Jakarta Batch model. Processing Each Customer The next artifact implements Jakarta Batch’s ItemProcessor interface. Java @Named("customerTierProcessor") @Dependent public class CustomerTierProcessor implements ItemProcessor { @Inject @BatchProperty( name = CustomerSegmentationPolicy.JOB_PARAMETER) private String thresholdsJson; private CustomerSegmentationPolicy policy; @PostConstruct void initialize() { policy = CustomerSegmentationPolicy.fromJson( thresholdsJson); } @Override public Customer processItem(Object item) { if (!(item instanceof Customer customer)) { throw new IllegalArgumentException( "Expected a Customer item"); } CustomerTier calculatedTier = policy.tierFor(customer.getTotalSpent()); if (calculatedTier == customer.getTier()) { return null; } return Customer.builder() .id(customer.getId()) .name(customer.getName()) .totalSpent(customer.getTotalSpent()) .tier(calculatedTier) .build(); } } Here the Jakarta Batch contract is explicit: ItemProcessor defines processItem(). The runtime calls that method for every item produced by the reader. The processor applies the segmentation rule and either returns the transformed customer or null. Returning null has a specific meaning in Jakarta Batch: the item is filtered and does not continue to the writer. The @BatchProperty is also part of the Batch integration. It receives the thresholds property defined for this job execution, allowing the processor to reconstruct the CustomerSegmentationPolicy before processing begins. Writing the Results The final artifact extends AbstractItemWriter, Jakarta Batch’s base implementation for writing a chunk. Java @Named("customerItemWriter") @Dependent public class CustomerItemWriter extends AbstractItemWriter { @Inject private CustomerRepository customerRepository; @Override public void writeItems(List<Object> items) { List<Customer> customers = items.stream() .map(this::toCustomer) .toList(); customerRepository.saveAll(customers); } private Customer toCustomer(Object item) { if (item instanceof Customer customer) { return customer; } throw new IllegalArgumentException( "Expected a Customer item"); } } writeItems() is defined by the Jakarta Batch writer contract inherited from AbstractItemWriter. Unlike the processor, which receives one item at a time, the writer receives a collection of processed items. In this case, the collection contains only customers whose classification changed, as the processor has already filtered the others. At this point, the Java components of the pipeline are as follows: Plain Text CustomerItemReader extends AbstractItemReader ↓ CustomerTierProcessor implements ItemProcessor ↓ CustomerItemWriter extends AbstractItemWriter These types are what connect the application code to the Jakarta Batch runtime. Connecting the Artifacts With JSL The Java classes define the behavior, but Jakarta Batch requires explicit mapping of the reader, processor, and writer to each job. This orchestration is described in JSL: XML <?xml version="1.0" encoding="UTF-8"?> <job id="customer-segmentation" xmlns="https://jakarta.ee/xml/ns/jakartaee" version="2.0"> <step id="recalculate-customer-tiers"> <chunk item-count="20"> <reader ref="customerItemReader"/> <processor ref="customerTierProcessor"> <properties> <property name="thresholds" value="#{jobParameters['thresholds']}"/> </properties> </processor> <writer ref="customerItemWriter"/> </chunk> </step> </job> The ref values correspond directly to the names declared with @Named in the Java classes: Java @Named("customerItemReader") @Named("customerTierProcessor") @Named("customerItemWriter") The XML therefore tells the Jakarta Batch runtime: for this step, use this reader, then this processor, and finally this writer. It also maps the thresholds job parameter into the processor property. The item-count="20" sets the chunk size for this sample. Jakarta Batch coordinates reading and processing, periodically invoking the writer according to the chunk lifecycle and establishing transaction and checkpoint boundaries. The value 20 is for demonstration; real applications should tune chunk size based on processing cost, database behavior, transaction size, throughput, and recovery requirements. This structure is recommended for the article: present the class declaration first, then describe the lifecycle methods inherited from or required by Jakarta Batch. This approach helps the sample teach the API rather than simply presenting isolated methods. Conclusion Jakarta Batch is valuable because it transforms large-scale data processing into a structured execution model, eliminating the need for custom loops and ad hoc background logic. By separating reading, processing, and writing, and introducing runtime concepts such as jobs, steps, checkpoints, restartability, and chunk-oriented execution, it provides enterprise applications with a predictable approach to handling workloads involving thousands or millions of records. This allows implementations to focus on business logic, while the Batch runtime manages repetitive execution concerns, making the model easier to understand, optimize, and scale as workloads increase.

By Otavio Santana DZone Core CORE
The Warning That Never Stops the Agent
The Warning That Never Stops the Agent

I went looking for a flag that could cap the amount spent on a session on DeepAgents. The kind of hard spending cap that stops an agent loop before it burns through a budget. What exists instead is a warning, and the gap between "warning" and "stopping" turned out to be the more interesting story. What Already Exists: A Number With No Teeth deepagents-code has a CostTrackingMiddleware that owns a thread's cumulative spend. This is a real, checkpointed dollar figure priced from actual per-request token usage, not an estimate. A separate feature built on top of that number, merged earlier, adds a one-time warning once the total crosses a configured threshold (default $50): Python # app.py, roughly what ships today threshold = self._session_cost_warning_threshold_usd if ( not self._session_cost_warning_shown and 0 < threshold < self._session_cost_usd ): self._session_cost_warning_shown = True self.notify( f"Estimated session cost is {format_cost(self._session_cost_usd)}, " f"above the configured {format_cost(threshold)} threshold. Consider " "/offload to reduce context usage or /clear to start fresh.", title="Session cost warning", severity="warning", timeout=12, markup=False, ) Read that once more: It's a self.notify(...) call. A toast! The agent's own loop has no idea this happened. Nothing in the code above touches the graph, the model call, or the next tool invocation. If you're watching the terminal, you see the warning and can intervene by hand (/offload, /clear, or just killing the process). If you're not watching and it's a headless CI run — for example, an unattended overnight session or a tool-call loop that's quietly retrying against a flaky API — this number keeps climbing, and nothing stops it. That's the gap: a cost number that can only ever inform a human, never the loop that's actually spending the money. Why the Fix Isn't "Just Check the Number Somewhere" The obvious instinct is to add a if cost > limit: stop check. The real work is in where that check has to live and what "stop" has to mean to a LangGraph agent loop. Where: The check needs to run before the next model call, not after. Checking after a call has already happened is too late to prevent its cost. LangGraph gives middleware a before_model hook for exactly this. LangChain's own ModelCallLimitMiddleware (a call-count limiter, not a cost one) already establishes the pattern: check a condition in before_model, and if it's tripped, return an update that redirects the graph to end instead of letting the model call happen. What "stop" means: A plain return None from before_model just lets the loop continue. To actually halt, the hook needs @hook_config(can_jump_to=["end"]) and has to return {"jump_to": "end", ...} which is a real graph-control signal, not a value the caller has to notice and act on: Python # cost_tracking.py — the actual hook, as committed @hook_config(can_jump_to=["end"]) def before_model( self, state: CostState, runtime: Runtime[ContextT], ) -> dict[str, Any] | None: """Halt the run before the next model call if the hard cost cap is met. Checked against the checkpointed cumulative total from the *previous* step -- the same figure the TUI's soft warning reads via `_set_session_cost` -- so this fires at the same point in the loop a user would already have seen the warning toast, just before the next request that would push spend further over the configured cap. """ if self._nested or self._hard_limit_usd is None: return None total_usd = state.get("_session_cost_usd") if ( isinstance(total_usd, bool) or not isinstance(total_usd, int | float) or not math.isfinite(total_usd) ): return None if total_usd < self._hard_limit_usd: return None if self._exit_behavior == "error": raise CostLimitExceededError(total_usd, self._hard_limit_usd) limit_message = _build_cost_limit_message(total_usd, self._hard_limit_usd) return {"jump_to": "end", "messages": [AIMessage(content=limit_message)]} Two details worth calling out, because they're the kind of thing that looks like overcaution until you hit the failure it's guarding against: isinstance(total_usd, bool) before the numeric check. In Python, bool is a subclass of int, so isinstance(True, int | float) is True and True < 5.0 evaluates fine (True == 1). Without the explicit bool guard, a stray True sitting in a state field meant for a float would silently be treated as $1.00 and could either falsely trip the halt or falsely pass it, depending on the cap. Cheap to guard against, expensive to debug if you don't. Checked only on the non-nested instance. CostTrackingMiddleware also runs on subagents, where its _session_cost_usd channel tracks that subagent's own local spend before it's transferred back to the parent's running total. Checking the hard cap there would be checking the wrong number - a subagent's small local total against a cap meant to bound the whole session's spend. self._nested gates this out entirely. The Part That Isn't the Algorithm: Getting the Number Across a Process Boundary Here's what made this bigger than a one-file change: dcode doesn't run the agent loop in the same process as the CLI. It starts a LangGraph server in a subprocess and talks to it over langgraph-sdk. Every configuration value the agent needs, like model name, sandbox type, recursion limit, and now this cap, has to survive that boundary, which in this codebase means round-tripping through environment variables on a ServerConfig dataclass: Python # _server_config.py max_cost_usd: float | None = None """Explicit hard cap, in USD, on the main thread's cumulative estimated cost.""" def to_env(self) -> dict[str, str | None]: return { ... "MAX_COST_USD": ( str(self.max_cost_usd) if self.max_cost_usd is not None else None ), ... } @classmethod def from_env(cls) -> ServerConfig: return cls( ... max_cost_usd=_read_env_float("MAX_COST_USD", default=None), ... ) _read_env_float didn't exist before this - every other numeric config value in this file is an int (recursion_limit, turn counts), so there was no float-reading helper to reuse. One new function, matching the existing _read_env_int's shape exactly, and the round-trip works the same way every other config value already does. The full path, in order: a --max-cost CLI flag → a resolver that checks the flag, then a config.toml entry, then "disabled" → create_cli_agent(max_cost_usd=...) → ServerConfig → serialized to an environment variable → the subprocess reads it back → create_cli_agent again, this time inside the subprocess → CostTrackingMiddleware(hard_limit_usd=...). Six hops for one float, and every one of them was necessary - skip the ServerConfig round-trip and the flag works in a unit test but silently does nothing the moment you actually run dcode, because the subprocess that runs the real agent loop never sees it. Verifying It Live Unit tests covering the before_model logic in isolation are necessary but not sufficient here. They'd pass even if one of those six hops silently dropped the value, because a unit test calls the middleware directly and never exercises the subprocess boundary at all. The only way to know the flag actually works is to run the real CLI against a real model: Shell $ dcode -n "Read sample.txt, then read it again, then read it a third time. \ Do this as three separate tool calls, one per turn." \ --model anthropic:claude-haiku-4-5 \ --max-cost 0.0001 \ --max-turns 6 I'll read sample.txt, then read it again on the next turn, then a third time on the turn after that. First read: Calling tool: read_file Session halted: estimated cost $0.02 has reached the configured limit of <$0.01. Raise the limit (e.g. `--max-cost`) or start a new session to continue. Task completed Usage Stats Provider Model Reqs InputTok OutputTok Cost anthropic claude-haiku-4-5 1 13.6K 147 $0.02 The model was told to make three tool calls, one per turn. It made exactly one. The middleware checkpointed that turn's real cost (0.02), the next beforemodel check found the cumulative total over the(deliberately absurd) 0.0001 cap, and the graph jumped to end with the injected message instead of continuing to spend on turns two and three. Total cost of proving this worked: two cents. Why This Is Worth a Hard Stop and Not Just a Bigger Warning You could imagine closing this gap by making the warning louder: repeat it every turn instead of once, or block user input until it's acknowledged. That doesn't fix the actual failure mode, which is specifically the unattended case. A louder toast is still a toast. Nothing short of a return value the graph itself has to obey closes that gap, which is why the fix has to live in before_model, not in the terminal UI layer where the existing warning already sits. It's also worth being honest about what this doesn't fix: the check runs before a model call, using the cost checkpointed from the previous one. A single turn can still overshoot the cap if the cap is 5.00 and the agent is at 4.99. The next call still happens in full and might land at 6.00 before the halt fires on the turn after. That's not a bug so much as an inherent property of checking after the fact rather than metering mid-request, and it's the same tradeoff that ModelCallLimitMiddleware makes for call counts. A cap is a backstop against runaway, unattended spend. It's not a precise billing guarantee down to the last cent. Takeaways The general lesson here isn't really about cost. It's that a warning and a limit are two different features wearing the same clothing, and it's easy to ship the first while believing you've shipped the second. The warning reads the same number, uses the same word ("threshold"), and looks like it's doing the same job right up until someone isn't in the room to read it. The tell is always the same: does the check return a value the system has to act on, or does it just call something with "notify" in the name? Second, a fix that only works in-process is only half a fix once your architecture has a subprocess boundary in it. The six-hop threading here wasn't extra caution, but it was the actual scope of the problem, and skipping any one hop would have shipped a flag that silently does nothing. Third, if you can run the real thing end-to-end for two cents, there's no good reason to trust a mock's word for whether a fix actually works.

By Ninaad Rao DZone Core CORE
Jakarta Faces Flow Scope: Managing Multi-Step UX Without Session State
Jakarta Faces Flow Scope: Managing Multi-Step UX Without Session State

Multi-step flows are common in UX, including onboarding, checkout, account setup, approval processes, configuration wizards, and administrative tasks. These require users to move through multiple screens while continuing a consistent working state. The challenge is to keep this state active for the duration of the interaction, but not beyond. Request scope is too short, while session scope often extends longer than the business process needs. Jakarta Faces handles this with @FlowScoped, which manages state based on the lifecycle of a flow instead of a single page or the entire session. This article uses a customer segmentation application to demonstrate how a flow can guide users through configuration, preview, and confirmation, while maintaining state across each step. This approach creates a cleaner model for wizard-style UX: the scope begins when the user enters the flow, persists during navigation, and ends upon exit. Why Jakarta Faces Still Matters Jakarta Faces continues to be relevant because many enterprise applications are developed and maintained by teams with strong Java expertise. In these environments, a server-side UI framework limits context switching, keeps validation and navigation close to the application model, and allows teams to reuse the same language, dependency injection model, and enterprise APIs throughout the stack. Architecturally, if the team is proficient in Java and the application is form-driven, workflow-oriented, or back-office focused, introducing a separate frontend stack does not necessarily offer an advantage. Component libraries such as PrimeFaces further support this approach. Rather than building tables, dialogs, forms, charts, wizards, and validation from scratch, teams can use reusable components within the Jakarta EE programming model. This can accelerate delivery and lessen the need for custom frontend infrastructure. While the decision should be based on product and team context, for Java-focused enterprise teams, Jakarta Faces is a pragmatic architectural choice, not just a legacy option. A few publicly documented examples of organizations that have used Jakarta EE/JSF or PrimeFaces include: NASAWalmart LabsLufthansaRakutenCommerzbankUnited NationsPenn State UniversityHarvard UniversityTelefonicaBig LotsComfortel (telecommunications)Various commercial banks and financial institutions Overview of Jakarta Faces Scopes Jakarta Faces applications maintain managed bean state for varying durations. Choosing the right scope is an architectural decision and must match the user interaction's lifetime. Some state is limited to a single HTTP request, a single page, a multi-step flow, or the entire user session. @RequestScoped is the shortest-lived scope. It creates a bean instance for each HTTP request and discards it when the request completes. It suits stateless actions, simple submissions, and operations that don't need to continue across navigation or Ajax interactions.@ViewScoped retains the bean while the user stays on the same Faces view. It is ideal for pages with forms, tables, filtering, pagination, dialogs, or Ajax interactions that update the same page multiple times. The state is discarded when the user navigates to a different view.@FlowScoped is intended for business interactions spanning multiple views, such as checkout, onboarding, configuration wizards, approval workflows, or customer segmentation. The bean remains active throughout the flow and is destroyed when the flow ends. Its lifecycle falls between view scope and session scope.@SessionScoped maintains state for the entire user session. It suits information that remains across multiple pages, such as user preferences or session-level context. Avoid using it for temporary workflow state, as this can unnecessarily extend the state’s lifetime.@ApplicationScoped has the broadest lifetime, sharing a single bean instance across the entire application and all users. It suits shared services, caches, configuration, or application-wide state, but not per-user or per-flow data unless explicitly designed for sharing and thread safety. Building a Multi-Step Experience With @FlowScoped This article focuses on Jakarta Faces flow, which models user engagements spanning multiple pages but shorter than a full HTTP session. In the customer-segmentation example, the administrator configures thresholds, previews their impact, reviews the configuration, and starts the operation. These steps form a single business process and should share the same state. Jakarta Faces represents this process with a flow definition and a flow-scoped managed bean. In this project, the flow resides in the customer-segmentation directory: reStructuredText src/main/webapp/ └── customer-segmentation/ ├── customer-segmentation-flow.xml ├── configure.xhtml ├── preview.xhtml └── review.xhtml Aligning the directory, flow identifier, and bean name clarifies their relationship. Here, the flow stays named customer-segmentation, defined in customer-segmentation/customer-segmentation-flow.xml, and the Java bean uses @FlowScoped("customer-segmentation"). The value passed to @FlowScoped ties the bean’s lifecycle to the corresponding Faces flow. The XML file defines the flow’s structure and navigation boundaries, but does not store business state. In this example, configure is the starting point, followed by preview and review. Each view has an identifier and references its corresponding XHTML document: XML <flow-definition id="customer-segmentation"> <start-node>configure</start-node> <view id="configure"> <vdl-document> /customer-segmentation/configure.xhtml </vdl-document> </view> <view id="preview"> <vdl-document> /customer-segmentation/preview.xhtml </vdl-document> </view> <view id="review"> <vdl-document> /customer-segmentation/review.xhtml </vdl-document> </view> <flow-return id="home"> <from-outcome>/index?faces-redirect=true</from-outcome> </flow-return> </flow-definition> The start-node specifies where the interaction begins. The <view> elements define the flow’s pages, and <flow-return> determines how the application exits. Returning the outcome home ends the flow, redirects the user to the dashboard, and discards the flow-scoped state. Navigation within the flow preserves the state, while exiting ends the conversation. On the Java side, CustomerSegmentationFlow manages the conversation state as both a named Faces bean and a flow-scoped bean: Java @Named @FlowScoped("customer-segmentation") public class CustomerSegmentationFlow implements Serializable { @Inject private CustomerSegmentationFlowService flowService; private CustomerSegmentationFlowState state; @PostConstruct public void initialize() { state = flowService.initializeState(); } // ... } @Named makes the bean accessible to Faces pages through Expression Language, while @FlowScoped("customer-segmentation") assigns its lifecycle to the specific flow. The state created during @PostConstruct remains across requests and page transitions during the flow, so CustomerSegmentationFlowState remains available throughout the interaction. This distinction sets @FlowScoped apart from @ViewScoped. With view scope, moving from configure.xhtml to preview.xhtml starts a new conversation. With flow scope, both views remain part of the same business interaction. The scope follows the conversation, not individual pages. Navigation methods on the bean then return outcomes that correspond to nodes defined by the flow: Java public String preview() { flowService.preview(state); return "preview"; } public String review() { return "review"; } Returning "preview" moves the user to the preview view in the flow definition, while "review" advances to the next view. Since both destinations are within customer-segmentation, the same flow-scoped bean and its state remain active. The final action demonstrates the other side of the lifecycle: Java public String execute() { long executionId = flowService.start(state); // message handling omitted return "home"; } Home is not a page within the wizard. It corresponds to the <flow-return id="home"> element in the XML definition. When this outcome occurs, Faces exits the flow, redirects to the dashboard, and discards the associated state. @FlowScoped is ideal for processes such as checkout, onboarding, registration, approval, configuration, and administrative wizards. It provides a scope broader than a single page but narrower than a session. Instead of using @SessionScoped for temporary workflow data, the state exists only for the workflow's duration. This article covers the Faces flow: how pages are connected, how state persists between them, and how entering and exiting the flow manages the Java bean’s lifecycle. The full application also includes MongoDB persistence, dashboard services, preview calculations, and Jakarta Batch processing. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. Here, the emphasis stays on the user interaction represented by Configure → Preview → Review → Exit. Conclusion @FlowScoped provides Jakarta Faces with an effective way to model multi-step user experiences as a single business conversation. Rather than storing temporary workflow state in @SessionScoped or reconstructing it between views, the flow maintains state only while the user progresses through the defined steps and releases it when the flow concludes. For Java-focused enterprise teams, this approach simplifies wizards, onboarding, approvals, checkout flows, and configuration processes by keeping navigation, state, and lifecycle consistent with the user experience.

By Otavio Santana DZone Core CORE
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do

For the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening. The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified. This is a significant architectural shift. The emerging pattern looks less like: Plain Text prompt → LLM → source code → trust it (the infamous vibe ding pattern) and increasingly like: Plain Text intent ↓ LLM ↓ restricted set of operations ↓ deterministic tools and validation ↓ result Several recent developments in mainstream developer tooling point in exactly this direction. Permission Is Becoming Separate From Intelligence On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations. Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals. That distinction is more profound than it may initially appear. The question is no longer merely: Can the agent perform this operation? It is: Is this agent authorized to perform this operation in this environment? Capability and authority are different things. A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so. A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities. Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering. The Model Is Becoming Replaceable Another development makes the separation even clearer. GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation. This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go. Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear. Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is: Plain Text stable environment stable tools stable constraints stable validation ↑ interchangeable models stable environment stable tools stable constraints stable validation ↑ interchangeable models The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system. Agents Are Separating From Editors VS Code's Agent Host architecture points in another related direction. The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts. Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code: Plain Text human ↓ editor ↓ language server ↓ compiler Agentic development introduces another participant: Plain Text human intent ↓ agent ↓ semantic tools ↓ compiler / tests / environment The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important. A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things. Verification Is Moving From Opinion to Execution A fourth development may ultimately be the most important. GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct. This should sound obvious. We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense. A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system. A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes: Plain Text generate ↓ compile ↓ test ↓ inspect diagnostics ↓ repair ↓ repeat That is substantially more robust than asking a model to inspect its own output and say whether it looks right. The role of the LLM is reasoning. The role of deterministic software remains enforcement. The Interesting Convergence These developments come from different parts of the development stack, but they point toward the same decomposition. An AI programming environment increasingly contains at least four distinct elements: A model that reasons and generatesA vocabulary of operations available to itA capability policy defining which operations it may useDeterministic mechanisms that decide whether the result is valid None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design. What If We Put the Boundary Into the Language? Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted. There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes? This is the idea I have been exploring with an open-source project called BUBAS. BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment. There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary. An order-processing application could, for example, expose operations such as: Plain Text LOAD_ORDER ORDER_TOTAL CUSTOMER_RISK APPROVE REJECT REQUEST_APPROVAL An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed. If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that. Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem. This Is Not a Sandbox Make the distinction carefully. A restricted language does not magically make its host application safe. If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong. The useful guarantee is narrower: Generated business logic can only name operations that the application deliberately made part of its vocabulary. That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment. These operate at different semantic levels. Domain Capabilities Are More Interesting Than Operating-System Capabilities Operating-system permissions are necessary, but business applications eventually need a richer vocabulary. Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful. It still does not answer questions such as: May this program approve an order?May it request approval but not approve directly?Can it read customer risk information?Can it initiate a payment?Can it calculate a premium but not change the underlying policy? Those are domain capabilities. General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction. A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations. We May Be Seeing the New Shape of the AI Programming Stack None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software. But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change. The industry is gradually replacing: Trust the model. with: Give the model precisely defined capabilities and verify what it produces. That is a much more promising engineering principle. BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.

By Peter Verhas DZone Core CORE
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.

By Akhil Madineni DZone Core CORE

Culture and Methodologies

Agile

Career Development

Methodologies

Team Management

Can Your Team Name the Work It Already Runs With AI?

September 25, 2026 by Stefan Wolpers DZone Core CORE

One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph

September 25, 2026 by Akhil Madineni DZone Core CORE

How to Build an Asynchronous AI-Content Review Workflow in C#

September 23, 2026 by Brian O'Neill DZone Core CORE

Data Engineering

AI/ML

Big Data

Databases

IoT

Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

September 30, 2026 by Ai Cerrudo

Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery

September 30, 2026 by Uthej Mopathi DZone Core CORE

OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay

September 29, 2026 by DZone Staff

Software Design and Architecture

Cloud Architecture

Integration

Microservices

Performance

Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus

September 30, 2026 by Akhil Madineni DZone Core CORE

Predict, Repeat, Improve: Deterministic Simulation Testing Explained

September 29, 2026 by Ammar Husain DZone Core CORE

Detection and Response Did Its Job. Now Someone Has to Actually Fix It.

September 29, 2026 by Philip Piletic DZone Core CORE

Coding

Frameworks

Java

JavaScript

Languages

Tools

Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery

September 30, 2026 by Uthej Mopathi DZone Core CORE

A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data

September 29, 2026 by Jubin Soni, FBCS DZone Core CORE

Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads

September 29, 2026 by Otavio Santana DZone Core CORE

Testing, Deployment, and Maintenance

Deployment

DevOps and CI/CD

Maintenance

Monitoring and Observability

Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus

September 30, 2026 by Akhil Madineni DZone Core CORE

Predict, Repeat, Improve: Deterministic Simulation Testing Explained

September 29, 2026 by Ammar Husain DZone Core CORE

A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data

September 29, 2026 by Jubin Soni, FBCS DZone Core CORE

Popular

AI/ML

Java

JavaScript

Open Source

Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

September 30, 2026 by Ai Cerrudo

Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery

September 30, 2026 by Uthej Mopathi DZone Core CORE

OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay

September 29, 2026 by DZone Staff

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×