The Testing, Tools, and Frameworks Zone encapsulates one of the final stages of the SDLC as it ensures that your application and/or environment is ready for deployment. From walking you through the tools and frameworks tailored to your specific development needs to leveraging testing practices to evaluate and verify that your product or application does what it is required to do, this Zone covers everything you need to set yourself up for success.
Metamorphic Testing For LLMs: The Oracle Problem's Most Underused Answer
From Issue to Reviewable Patch: Build a Sandboxed Coding Agent With Deep Agents and Automated Tests
Throughout my career, I've watched sprints stall in the same place. Everything was going fine: the features were merged, the requirements were clear. But then we hit a roadblock: waiting on test cases. Because the QA team was hand-converting Jira stories into steps and expected results. Most test automation strategy conversations skip right past this authoring bottleneck. I've done the conversion work myself. It's not hard work, but it's necessary, slow, and repetitive. Worst of all, it's disconnected from the actual skills that made anyone want to work in quality engineering in the first place. In my previous article, I drew a line between two architectures that both get sold as "AI test generation." One is simply a large language model behind a prompt template. The other is an actual agent that reads your requirements, your attachments, and your existing test libraries before it writes anything. Agentic test creation means an AI agent reads your requirements, attachments, and existing test library, then drafts structured test cases traced to each acceptance criterion, with a human review gate before anything enters the suite. That article included a 6-step example of what an agentic pipeline does with a promo code user story. The result: 7 cases reused, 6 generated, 1 of which was rejected at review. This time I want to expand that example into a full walkthrough, drawn from my own experience. The example isn't special. These are the same plain-language requirements your team probably already writes. But the process and outcome are interesting. Starting With a Plain-Language Requirement When I say plain-language requirements, I mean the artifacts an agile team already produces. User stories. Acceptance criteria, whether they're bullet points or structured Gherkin scenarios. The Jira ticket itself, with its comments and revision history. Even the wireframe or screenshot somebody attached during refinement. None of it is specifically written for an AI. But the good news is that none of it needs to be. That's the premise behind AI test case generation from user stories: the story is the spec, and its criteria are the test conditions. It's also what shift-left testing looks like in practice. If the requirement itself is the test input, test design starts the moment the story is written. The Agent Loop vs. an LLM Call I covered the agent loop in detail last time, so I'll just give an overview here. A typical single LLM call is a stateless function. You call, you get an answer. You move on to the next call. But an agent is a loop: it reasons about a goal, calls tools to gather information, observes the results, and revises its plan before producing output. That's the reasoning-and-acting loop the ReAct paper formalized (Anthropic's Building Effective AI Agents is my recommended primer). Applied to test creation, the agent loop parses the story and its criteria, pulls the linked test cases, reads the attachments, and builds a coverage map — all before generating any steps. The Agentic Test Case Walkthrough Here is our requirement: As a returning customer, I want to apply a promo code at checkout so that my discount is reflected in the order total. The acceptance criteria attached to the story (let's call it PROMO-214): A valid promo code reduces the order total and shows a discount line item.An invalid or expired code shows an inline error and leaves the total unchanged.A promo code and a gift card can be applied to the same order; the promo discount is applied first.Percentage discounts round half up to the nearest cent at the order level, not per item.Removing a code restores the original total. Plus there's an attachment: checkout-mockup.png, showing the promo field, the discount line, and the gift card entry point. That's our entire input package. Here's what it looks like: What the Agent Surfaces Before Generating Before writing anything, the agent first checks the test cases linked to the checkout area. And it finds 40! It maps the story's scenarios against them and proposes 7 for reuse: valid code, invalid code, expired code, empty field, case sensitivity, code removal, and re-application after removal. In a generic tool, those 7 would come back as duplicates. Here they arrive as reuse suggestions with their existing case IDs, so their regression history stays attached. The Output Our agent now generates 6 new cases against actual gaps. Here's one: TC-1207 · Apply a promo code and a gift card to the same order (AI-generated · awaiting review) Preconditions Returning customer signed in. Cart contains 2 items ($39.97 and $40.02) totaling $79.99. Valid 20% code SAVE20. Gift card balance of $30.00 on the account. Saved card on file. Step 1 Proceed to checkout. Expected: order summary shows $79.99. Step 2 Enter SAVE20 and select Apply. Expected: discount line shows −$16.00; total updates to $63.99. Step 3 Apply the gift card. Expected: gift card line shows −$30.00; total updates to $33.99. Step 4 Place the order. Expected: confirmation lists both adjustments; $33.99 is charged to the saved card; remaining gift card balance is $0.00. Traceability PROMO-214, acceptance criteria 1, 3, 4 (primary: 3) Let's look at what the agent did with acceptance criterion 3. The criterion states an ordering rule: promo first, then gift card. The generated steps verify the amounts in that order, with the expected results as the actual arithmetic. That arithmetic came from reading the criteria together: 20% off $79.99 is $15.998, which becomes $16.00 because criterion 4 rounds at the order level. Rounded per item, the same discount would be $15.99. The rate, the ordering rule, and the rounding rule all interact. A second generated case, TC-1208, targets the expired-code path in a checkout that already has a saved payment method on file. It renders like this: The remaining 4 cases cover rounding at the half-cent boundary, a gift card that exceeds the discounted total, code removal after a gift card is applied, and discount persistence across a session timeout. None of these are uncommon, but they're typical cases we might skip when we run out of time in a sprint. Just as telling is what the agent didn't generate. Nothing in the batch covers 2 browser tabs applying codes to the same cart, and nothing checks whether the inline error is announced to a screen reader ... because no criterion mentions either. The agent's coverage tracks what's written down—and when it strays past that, the review gate catches it. Deciding what should have been written down is still your job. One objection you might raise: these are structured test cases, not automated end-to-end scripts. That's true. But the structure is the point. A case with discrete steps and explicit expected results is an automatable artifact. Whether a person executes TC-1207, an automation engineer scripts it, or an execution agent picks it up downstream, the hard part is already done: what to verify, in what order, and with what data. Prose test ideas can't make that handoff. But steps with expected results can. Your Output Is Only as Good as Your Input Next, let's take a look at the limits of this approach. Like most solutions, agentic test cases have some best practices that should be followed. First, vague acceptance criteria produce vague steps. If criterion 4 hadn't specified order-level rounding, the agent would have had to infer a rounding rule, and a reviewer would have needed to confirm it. What Does a Testable Criterion Look Like? So what does a testable criterion look like? Compare "discounts should work with gift cards" to criterion 3 above: a promo code and a gift card can be applied to the same order, and the promo discount is applied first. The first version tells the agent a feature exists. The second gives it an ordering rule it can verify with arithmetic. A few habits close this gap: give ordering rules explicitly, specify rounding and limits, name the expected error behavior, and attach your mockups. None of this is new, right? Testers have been pushing for the best practices for years. The agent just makes it more important. Second, a thin or messy test library weakens the reuse step. An agent can only propose reusing cases it can find! And finally, treat the review gate as mandatory. In the original run of this example, our reviewer edited 2 cases and rejected 1. In my previous article, I recommended tracking the reviewer rejection rate; I'll refine that here: track 2 numbers separately, the rejection rate and the edit rate. Rejection rates going up usually means the context feeding the agent is broken. Maybe it's a thin test library, or stories that aren't linked to their cases, or even criteria that are silent on a whole scenario. Edit counts going up usually means your acceptance criteria are vague. Rejected Test Cases Let's pause for a moment to look back at our rejected case from earlier. We didn't dig into that much. Agentic failures sometimes don't look like failures. In this example, our reviewer rejected a session-timeout case. It had clean steps, specific expected results, and was formatted exactly like the other 5. But the problem was that we gave no acceptance criterion on what happens to a discount when a session expires. The agent inferred a behavior and then tested its own inference. That's the pattern you should use to train reviewers: the most dangerous output isn't the sloppy case; it's the really good-looking case... that just happens to verify a requirement that no one wrote. The good news is that the rejection itself was useful. It made its way back to the product owner as a requirements question, which is precisely the kind of gap that previously wouldn't surface until production. Implementing Agentic Test Creation Finally, let's look at options for implementing agentic test creation. As is often the case, you can build, or you can buy. Buy. Vendors have started building true agentic solutions. For example, Agentic Test Creation in Tricentis qTest is one implementation of this pattern. It runs the loop inside the test management platform itself, so the reuse suggestions and the review gate land where the test cases already live. Build. You can also build this yourself. The building blocks are all there: the ReAct loop, an LLM that calls your tools, and community-maintained open-source connectors like the MCP Atlassian server that lets an agent read Jira stories directly. A motivated platform team can assemble this pipeline themselves. The benefit to buying? The plumbing. Existing test libraries, traceability links, and review workflows are already wired together. This can matter more than you might think. Reuse detection is only as good as the agent's view of what exists, and a platform that already holds your test library already has that view. Same story for the review gate: it's a workflow with roles, permissions, and an audit trail, which is exactly the kind of thing that's boring to build and easy to underestimate. The benefit to building, of course, is flexibility. You pick the model, you own the prompts, you can encode your team's house style for test cases, and you can wire in internal systems no vendor will support. But be sure you budget honestly. Writing an agent that drafts steps from a story might be a weekend prototype. But the review workflow, the permissions, and keeping the prompts current as your requirements evolve are real work. That you now own. How to Evaluate Agentic Test Creation Tools Whichever route you take, I'd ask these 3 questions: Does the solution read everything attached to the requirement, including images?Does the solution propose reuse before it generates?Does every machine-written case pass through a review gate before entering the project? A test automation strategy is ultimately a decision about where humans create the most value. When structured test cases can be drafted from the requirements your team already writes, the hours that went into transcription can instead move to the work only humans can do: exploratory testing, risk analysis, and deciding what should be tested in the first place. Revisiting Your Test Automation Strategy My readers may recall my personal mission statement, which I feel can apply to any IT professional: "Focus your time on delivering features/functionality that extends the value of your intellectual property. Leverage frameworks, products, and services for everything else." — J. Vester Your team already writes user stories and acceptance criteria. Those artifacts are enough to drive end-to-end test creation, provided the system reading them knows what already exists. Transcribing them into test steps by hand never really added much value. But a context-aware agent drafted from them just might. Have a really great day!
Browser agents become useful when they can do more than reach the right page and trigger the right control. A click is only an attempted action; it is not proof that a business operation completed. A form can submit while validation fails, a checkout button can respond while the API returns an error, and a success-looking route can render stale state. Playwright MCP is well suited to closing that gap because it exposes browser interaction through structured accessibility snapshots and adds explicit testing tools for checking visible elements, text, lists, and form values. The result is an agent loop that can treat verification as a first-class phase rather than as an optimistic interpretation of the previous action. An Action Is Not a Result The central design rule is simple: every state-changing action should have a postcondition. In an ordinary browser automation script, success is often inferred from the absence of an exception. That standard is too weak for an autonomous agent. Playwright MCP interaction tools such as browser_click operate on element references taken from accessibility snapshots, and most actions return an updated snapshot after triggered browser work settles. That makes the post-action page state immediately available, but the controller still has to decide what evidence counts as success. A useful contract separates intent, action, and evidence. For a task such as submitting an order, the action is a click on the final submission control. The evidence might be a visible confirmation heading plus an order identifier. The agent should not report completion until those conditions are independently checked. With the testing capability enabled, Playwright MCP provides browser_verify_element_visible, browser_verify_text_visible, browser_verify_list_visible, and browser_verify_value. Verification calls return Done on success and an error on failure, which creates a clean boundary between a browser action and a verified outcome. A minimal controller can therefore make the verification step explicit: TypeScript await client.callTool({ name: "browser_click", arguments: { target: submitRef } }); await client.callTool({ name: "browser_wait_for", arguments: { text: "Order confirmed" } }); const proof = await client.callTool({ name: "browser_verify_element_visible", arguments: { role: "heading", accessibleName: "Order confirmed" } }); if (proof.isError) { throw new Error("Order submission could not be verified"); } The important property is not the specific wrapper around callTool; it is the control flow. Action and verification are different operations, and a failed verifier changes the task status from “completed” to “unconfirmed.” browser_wait_for is appropriate when a specific asynchronous transition must finish, while Playwright MCP already waits for triggered navigation and network activity after most actions. Fixed sleeps should remain a last resort because the server supports waiting for text to appear or disappear directly. Verification Should Match the Business Outcome A reliable verifier checks the state that matters to the task rather than a convenient visual change. A button becoming disabled proves only that the button changed. A toast saying “Saved” is stronger, but still may not prove persistence if the application updates optimistically. Browser-level evidence becomes stronger when multiple independent signals agree: semantic UI state, the resulting page structure, and relevant network activity. Playwright MCP exposes each of these forms of evidence through snapshots, verification tools, and network inspection. Playwright MCP exposes network inspection through browser_network_requests and browser_network_request, allowing an agent to locate a relevant request and inspect its details. Console messages are also available through the core browser_console_messages tool. Those channels are useful when a task appears successful in the DOM while a background request fails or the page emits an uncaught error. Network evidence should still be tied to application semantics; an HTTP response by itself does not establish that the intended record contains the correct data. For a profile update, the strongest browser-side check may be a round trip: submit the change, wait for the completion signal, navigate away or reload, then verify the field value from the newly rendered state. browser_verify_value supports textboxes, checkboxes, radios, comboboxes, and sliders, so persistence checks can stay semantic instead of scraping raw HTML. TypeScript await client.callTool({ name: "browser_verify_value", arguments: { type: "textbox", element: "Display name", target: displayNameRef, value: expectedName } }); This pattern is especially important for agentic workflows because planning logic can be probabilistic while verification can remain deterministic. The model may choose among several valid ways to reach a form, but the acceptance criterion can still be exact: a heading exists, a field equals an expected value, a list contains required entries, or confirmation text is visible. Playwright MCP’s verification tools are designed around those concrete browser states. Accessibility Snapshots Make the Feedback Loop Precise Playwright MCP uses accessibility snapshots as the primary representation for agent interaction. A snapshot contains roles, accessible names, text, and element references used by subsequent tool calls. This is materially different from relying on screenshots as the main control surface. Screenshots remain valuable for visual diagnostics, but the MCP documentation explicitly directs actions toward snapshots rather than screenshot coordinates. That distinction improves verification quality. A semantic check such as “heading named Order confirmed is visible” is less ambiguous than a model deciding whether a collection of pixels resembles a success page. It also aligns the agent’s evidence with the same roles and names used by Playwright locators. When only part of a large page matters, browser_find can search the accessibility snapshot and return matching nodes with local context, reducing the need to repeatedly consume the full tree. Screenshots still have a place when the requirement is inherently visual, such as confirming layout, clipping, or rendering. For correctness of transactional browser work, however, semantic evidence should dominate. Tracing can then provide failure forensics rather than primary success criteria. With the devtools capability, Playwright MCP can record traces containing DOM snapshots, screenshots, network activity, console logs, and timing, making an unverified or failed run reproducible after the fact. Successful Agent Runs Can Become Regression Tests Verification becomes more valuable when it survives beyond a single agent session. Playwright MCP’s testing capability records matching expect(...) code for verification tools, and action responses can include generated Playwright code. The documentation explicitly shows an exploratory sequence being assembled into a conventional Playwright test. That creates a productive path from autonomous exploration to deterministic regression coverage. A verified flow can therefore graduate into a compact test instead of remaining hidden inside an agent transcript: TypeScript test("submits an order", async ({ page }) => { await page.getByRole("button", { name: "Place order" }).click(); await expect( page.getByRole("heading", { name: "Order confirmed" }) ).toBeVisible(); await expect( page.getByText(expectedOrderNumber) ).toBeVisible(); }); The same principle should shape production configuration. Only required capabilities should be exposed, isolated sessions should be preferred for repeatable runs, and origin restrictions can reduce accidental navigation. Playwright MCP supports capability selection, isolated profiles, allowed and blocked origins, secrets redaction, and configurable timeouts. Its documentation also warns that origin controls and secret handling are convenience defenses rather than security boundaries, so client-level permissions remain necessary when an agent can perform consequential actions. Conclusion A browser agent becomes dependable only when completion means more than “the click happened.” Playwright MCP provides the pieces required for a verification-centered design: structured accessibility snapshots for precise targeting, explicit verification tools for semantic postconditions, waiting primitives for asynchronous transitions, network and console evidence for deeper diagnosis, and traces for failed-run analysis. The strongest implementation treats every consequential action as a hypothesis that must be proven by observable browser state. That shift turns browser automation from a sequence of hopeful interactions into a controlled execution loop whose results can be checked, explained, and eventually converted into durable Playwright regression tests.
Most AWS migration projects don't fail because of technical complexity. They fail because teams treat migration as a single activity rather than a set of distinct strategies applied to different workloads. AWS defines seven migration strategies — the 7Rs — that determine how each application moves to the cloud. The decision of which strategy applies to which workload has more impact on project cost, timeline, and outcome than any architectural choice you'll make after. Yet in practice, most teams default to "lift-and-shift everything" without evaluating whether that's appropriate. This article presents a practitioner's framework for classifying workloads into the 7Rs, based on delivering 50+ AWS migrations across fintech, SaaS, healthcare, and e-commerce. The 7R Strategies Retire Not every workload deserves migration. During discovery, you will invariably find applications that are redundant, unmaintained, or replaceable. In a typical enterprise portfolio of 20–40 applications, 10–20% qualify for retirement. Decision criteria: No active users, duplicate functionality already covered by another system, or maintenance cost exceeds business value. Common mistake: Teams skip this step because retiring applications requires stakeholder conversations. The result is migrating dead applications that consume compute budget indefinitely. Retain Some workloads shouldn't migrate in this wave. Applications with deep hardware dependencies, pending end-of-life within 12 months, or complex regulatory constraints that require legal review before cloud deployment are candidates for retention. Decision criteria: High migration complexity combined with low business urgency, or external constraints that prevent cloud deployment within the project timeline. Retain is not "never migrate." It's "not now." Document these workloads with a future migration path and trigger conditions. Rehost (Lift-and-Shift) Moving applications to EC2 or containers without code modifications. AWS Application Migration Service (MGN) automates this by continuously replicating servers and orchestrating cutover with minutes of downtime. Decision criteria: Application has a short remaining lifespan (1-2 years), speed of migration matters more than optimization, or the application is a black box with no available source code. Timeline: Days to weeks per workload. Trade-off: You gain cloud elasticity and pay-as-you-go pricing immediately, but you inherit all existing architectural inefficiencies. A poorly designed monolith on-premises becomes a poorly designed monolith on EC2. Relocate Hypervisor-level migration, primarily for VMware workloads moving to VMware Cloud on AWS. The OS, application, and configuration remain untouched. Decision criteria: Large VMware estate, tight data center exit deadline, and applications that cannot tolerate any configuration change. Replatform Migration with targeted adaptations to managed services. The application architecture stays intact, but you replace self-managed infrastructure components with AWS equivalents: Self-ManagedAWS ManagedOperational BenefitSelf-hosted PostgreSQLRDS for PostgreSQLAutomated backups, patching, failoverCron jobs on EC2EventBridge + LambdaNo server to maintain, pay-per-invocationSelf-managed RedisElastiCacheAutomatic failover, scalingNginx load balancerApplication Load BalancerManaged TLS termination, WAF integrationSelf-hosted ElasticsearchOpenSearch ServiceManaged cluster scaling, snapshots Decision criteria: Application is well-structured but operationally expensive. The team spends significant time on database maintenance, patching, backup verification, or scaling. Timeline: 2-4 weeks additional per workload compared to rehost. Trade-off: Moderate additional effort (schema compatibility testing, connection string changes) in exchange for a 40-60% reduction in ongoing operational cost. For most mid-complexity applications, replatforming represents the optimal balance between migration effort and long-term benefit. Refactor (Re-Architect) Rebuilding applications for cloud-native patterns: microservices decomposition, containerization (ECS/EKS), serverless (Lambda), event-driven architecture (EventBridge, SQS, SNS, Step Functions). Decision criteria: The application is a core business asset that needs capabilities the current architecture cannot deliver — true horizontal scaling, independent service deployments, multi-region active-active, or zero-downtime deployments. Timeline: Months. Budget accordingly. Trade-off: Highest upfront investment, but delivers the best long-term results in terms of deployment velocity, fault isolation, and scaling capability. Reserve this for 2-3 applications maximum in a migration portfolio. Repurchase Replacing custom-built software with a commercial SaaS product. The application doesn't move to AWS; it moves to a vendor. Decision criteria: The in-house application solves a problem that is not a core competency and commercially available alternatives have matured to cover your requirements. Common candidates: CRM, HR systems, monitoring, project management. The Decision Framework Classification should happen during the assessment phase, before any infrastructure work begins. For each workload, evaluate four dimensions: 1. Business Value How critical is this application to revenue generation or core operations? High: Core product, customer-facing, revenue-generatingMedium: Internal operations, supports core processesLow: Legacy, rarely used, or duplicate functionality 2. Technical Complexity How difficult is it to migrate given current architecture, dependencies, and state management? High: Stateful, tightly coupled, hardware dependencies, proprietary protocolsMedium: Standard web application with database, some external integrationsLow: Stateless, containerizable, standard protocols 3. Team Capacity Does your engineering team have the skills and bandwidth to support a complex migration approach? High capacity: Can support re-architecting alongside other workLimited capacity: Can handle replatforming with some external supportMinimal capacity: Rehost or retain is the only realistic option 4. Time Constraint How quickly must this workload be operational on AWS? Immediate (weeks): Data center exit, contract expiryStandard (1-3 months): Planned migration within a programFlexible (3-6 months): Can wait for deeper optimization Mapping Dimensions to Strategy Business ValueComplexityCapacityTimeRecommended StrategyLowAnyAnyAnyRetire or RepurchaseAnyHighLowImmediateRehost (with future replatform plan)MediumMediumMediumStandardReplatformHighMedium-HighHighFlexibleRefactorAnyAnyAnyBlockedRetain A Practical Example Consider a portfolio of 15 applications for a mid-size SaaS company: Plain Text ┌─────────────────────────────────────────────────────┐ │ RETIRE (3) │ │ - Legacy admin panel (replaced by new one 2024) │ │ - Internal wiki (moved to Confluence) │ │ - Prototype service (never went to production) │ ├─────────────────────────────────────────────────────┤ │ RETAIN (1) │ │ - Hardware security module integration │ │ (requires legal review for cloud deployment) │ ├─────────────────────────────────────────────────────┤ │ REHOST (4) │ │ - Backoffice tools (low traffic, stable) │ │ - Legacy reporting engine (EOL in 18 months) │ │ - Monitoring collector agents │ │ - Staging environment clone │ ├─────────────────────────────────────────────────────┤ │ REPLATFORM (5) │ │ - Main API (PostgreSQL → RDS, cron → Lambda) │ │ - Worker services (EC2 → ECS Fargate) │ │ - File processing pipeline (S3 + Lambda) │ │ - Authentication service (→ElastiCache for sessions│ │ - Notification service (→ SES + SQS) │ ├─────────────────────────────────────────────────────┤ │ REFACTOR (1) │ │ - Core product platform (monolith → microservices) │ ├─────────────────────────────────────────────────────┤ │ REPURCHASE (1) │ │ - Custom CRM (→ HubSpot) │ └─────────────────────────────────────────────────────┘ This distribution — 20% retire, 7% retain, 27% rehost, 33% replatform, 7% refactor, 7% repurchase — is representative of what I see in practice. The replatform bucket is almost always the largest. Migration Tooling Alignment Each strategy maps to specific AWS tooling: StrategyPrimary ToolsRehostAWS Application Migration Service (MGN), Migration HubReplatformDMS (databases), manual adaptation, Terraform/IaCRefactorECS/EKS, Lambda, Step Functions, custom developmentRelocateVMware Cloud on AWS AWS Migration Hub provides a unified tracking dashboard across all strategies. For database migrations specifically, AWS Database Migration Service (DMS) handles both homogeneous and heterogeneous migrations with continuous replication (CDC), enabling near-zero-downtime cutovers. Common Anti-Patterns "Rehost everything, optimize later." Teams that plan to rehost first and replatform in a second phase rarely execute phase two. The urgency disappears once applications are running, and the team moves to other priorities. If replatforming is the right strategy, do it during migration. "Refactor everything for cloud-native." The opposite extreme. Not every application needs microservices. A well-structured monolith running on ECS Fargate can serve thousands of requests per second with simpler operations than a distributed system. "One strategy for all workloads." Every application in the portfolio has different characteristics. The decision framework exists because one size does not fit all. Conclusion The 7R classification exercise takes 3-5 days for a typical portfolio. It requires involvement from engineering leads, product owners, and sometimes finance (for retire/repurchase decisions). The output - a workload-by-workload strategy map - becomes the foundation for accurate timeline estimates, resource planning, and budget allocation. Without it, you're building infrastructure for workloads that might not need to exist. For a comprehensive breakdown of migration costs, the full 6-phase delivery process, and AWS tooling details, see my complete AWS cloud migration guide.
It’s 2 AM. Your phone buzzes, the on-call alert flashes, and suddenly you are staring at a production outage that makes no sense. Following the logs, you get a hint: when you rerun the same scenario in staging, everything behaves perfectly. None of the quality gates/QA pipelines catch it, chaos experiments didn’t reproduce it, and now a ghost is chased that only appears when the system is under real-world pressure. Distributed systems are notorious for these “phantom failures” — rare timing-dependent bugs that surface unpredictably and vanish just as quickly. They are the kind of dreaded incidents that keep engineers awake at night because they are unreproducible. Take a real-world example: A service once crashed because two nodes tried to become leader at the exact same millisecond. In staging, the timing never aligned that way, so the bug remained invisible. But in production, under heavy load, it just happens, sending the system into chaos. Engineers spent days trying to recreate the failure, but without a deterministic replay, it's pure luck to get a reliable reproduction. Even when it happens, engineers may not be sure what caused it or how to reproduce it deterministically. Enter deterministic simulation testing (DST). DST builds a fully controlled, re-playable simulation of your system’s world — nodes, clients, clocks, network delays, partitions — so that even the most elusive bugs can be identified, captured, replayed, and studied. In this article, we will uncover how DST can transform those unpredictable 2 AM incidents into predictable, debuggable coordinates — giving you a new way to tame the chaos of distributed systems. Deterministic Simulation Testing Definition Deterministic simulation testing (DST) is a software testing methodology that places the system under test within a fully controlled, simulated environment. All sources of non-determinism — system clock, thread scheduling, network, disk — are intercepted and made deterministic. The key property is that, for a given initial seed and configuration, the entire execution is reproducible. The same sequence of events, faults, and outcomes will occur on every run with that seed. Key Concepts Breaking down the above definition, below are the key concepts for DST: Determinism → The system’s behavior is purely dependent upon its initial state and the seed. All non-deterministic sources are simulated to achieve determinism.Simulation → The system is run in a virtual environment that can simulate faults and control the passage of time. E.g., controlled clock skew introduced across various nodes, added network delays to achieve out-of-order event delivery.Reproducibility → Any failure or bug found during simulation can be reliably reproduced by rerunning the simulation with the same seed.Scenario Exploration → By varying the seed and/or simulation parameters, DST systematically explores a vast range of possible execution paths and failure scenarios. How DST Works Let's consider a simple scenario where two users update and read the same record in a very short interval. User 1 updates a record with a new value at time instance T0, and User 2 reads the same record at time instance T1. Note that the interval between T0 & T1 stays the same. In an ideal case (i.e., scenario 1), the new value is updated or written immediately, i.e., without any delay. Thus, when User 2 reads the same record at T1, it is able to read the latest value. In scenario 2 suppose the write is delayed due to network partitioning, disk write etc. User 2 thus sees the old value of the record even if it reads the value at the same time instance T1. Although a stale read may look trivial, it may lead to workflow halt, process crash, etc. in a complex real-world system. Imagine what could happen in a real-world system where: Multiple processes are scheduled for execution, within and across nodes.Multiple network calls are made between several nodes.Multiple operations are performed by several distributed processes on a single disk. Traditional testing strategies or frameworks are inherently constrained and thus can’t simulate such delays or faults. Because of this, it's nearly impossible to identify, catch, reproduce, or debug issues arising from such situations — rendering them unreliable or, at best, non-deterministic. To achieve determinism, the testing framework must take total control over the environment to intercept and manage all external interactions as described below: Controlled scheduling → Instead of relying on the operating system’s unpredictable thread/coroutine scheduler, the simulator provides its own deterministic scheduler. Thus ensuring various scheduling combinations are simulated.I/O mocking → All network calls, disk writes, and clock queries are routed through the simulator, allowing it to inject latency, drop packets, or change the time (e.g., clock skew).Single-threaded execution → Many DST frameworks run the entire distributed system stack within a single thread, completely stripping away the chaotic, unrepeatable nature of multi-threading. Thus, by eliminating real-world “flakiness,” DST allows developers to reproduce chaotic distributed system bugs with perfect precision, thanks to its inherent ability to replay any failing execution: Seed-based replay → The same seed reproduces the exact sequence of events, making debugging tractable.Time-travel debugging → Some platforms (e.g., Flashback) allow stepping backward and forward through execution, inspecting state at any point for a granular view of the system. DST Implementation Approaches and Architecture Patterns Below are two approaches for DST. Pluggable Non-Determinism Design the system so that all non-deterministic components (clocks, I/O, etc.) are pluggable. This strategy is used by TigerBeetle. This requires: Abstracting all system interactions behind interfaces.Providing both real and simulated implementations.Ensuring that the same codebase can run in both production and simulation by swapping implementations at startup. Pros: Deep control and minimal divergence between test and production code. Suitable for greenfield systems. Cons: Requires significant upfront design and is challenging to retrofit into existing systems. Deterministic Hypervisors and Emulation A more recent and flexible approach is to run unmodified binaries inside a deterministic hypervisor or emulation layer. This strategy is used by Hermit and Weave. Pros: Can test existing systems without code changes; language-agnostic; simulates the entire stack. Cons: May have performance overhead; some system behaviors may escape determinism if not fully intercepted. Benefits System employing DST benefits as below: Identify and reproduce rare failures → DST allows engineers to replay the exact sequence of events that led to a bug. This eliminates the frustration of “flaky” issues that appear inconsistently, making debugging far more reliable. Moreover, DST helps find bugs in execution paths unreachable by example-based tests.Improved developer productivity → Bugs are easier to reproduce, debug, and fix; less time spent on war rooms and emergency triage.Improved confidence in correctness → DST validates critical invariants (like consensus, failover, or transaction consistency) under controlled simulations. Engineers gain assurance that core distributed protocols behave as expected even under stress. Thus, preventing rare bugs from reaching production, increasing system uptime and user trust.Scalable debugging for complex systems → In microservice or event-driven architectures, DST helps tame the exponential growth of possible interleavings by focusing on deterministic seeds. This makes large-scale debugging more tractable. Challenges and Limitations While DST provides unparalleled confidence, it requires significant architectural investment. Retrofitting DST into existing systems may require significant refactoring. It can be highly intrusive, requiring developers to write custom code or frameworks, as production code often cannot rely on external third-party libraries that invoke un-mocked I/O or system calls. Moreover, DST requires careful modeling of external systems to avoid missing integration bugs. Ensuring sufficient coverage without combinatorial explosion is a major challenge — especially in modern systems with multiple integration points. Below are gaps in tooling that limit DST outcomes: Language and platform support → Not all languages and runtimes have mature DST frameworks.Hypervisor limitations → Deterministic hypervisors may not support all system calls or hardware features. DST Comparison and Applicability DST vs. Chaos Engineering DST is proactive and enables perfect reproducibility. It is best suited for development and pre-production, catching bugs before they reach users. It can simulate production chaos in minutes, and every failure is a permanent regression. Chaos engineering is reactive, non-deterministic, and validates the behavior of real deployments. It is essential for catching issues arising from real infrastructure, misconfigurations, or dependencies that simulation cannot model. However, it cannot guarantee coverage or reproducibility, and carries the risk of impacting users. In essence, both DST and chaos engineering are complementary to each other and are necessary for comprehensive reliability. DST in Functional vs. Performance Testing Functional Testing DST is ideally suited for functional testing. Validates correctness under all possible interleavings, failures, and workloads.Checks invariants, safety properties, and liveness under stress.Finds rare, timing-dependent bugs that are invisible to example-based tests. Performance Testing DST is not primarily designed for performance testing. The simulated environment does not reflect real hardware performance, network latency, or throughput.Time is virtualized and compressed; I/O is in-memory.Performance metrics (latency, throughput) measured in simulation may not correspond to real-world values. However, DST can be used to: Validate performance-related invariants (e.g., absence of deadlocks, progress under load).Simulate pathological scenarios (e.g., extreme contention, resource exhaustion) to observe system behavior. Recommendation Combine DST for functional correctness with real-world performance and benchmarking suites for comprehensive validation. DST Applicability to AI/ML and Agentic Systems AI/ML systems, especially those based on large language models (LLMs) and agentic workflows, are fundamentally non-deterministic. This makes traditional testing and debugging extremely difficult, with “heisenbugs” that vanish when observed. DST can be adapted to AI/ML systems by creating controlled, simulated environments for agents to operate in. Or using a hybrid approach of combining deterministic components (rule-based logic) with LLM-driven reasoning, using seeds to replay failures. Case Studies FoundationDB, with its deterministic simulator tool, achieved legendary reliability by running trillions of simulated CPU-hours, finding and fixing every known bug before production.TigerBeetle built a Viewstamped Operation Replication simulator (VOPR) to simulate financial transaction systems, catching subtle bugs in consensus and replication.Ethereum Merge used Antithesis to test the transition to Proof-of-Stake, simulating multiple client implementations in a deterministic environment. Conclusion Deterministic simulation testing (DST) represents a paradigm shift in the testing and validation of distributed systems. By enabling exhaustive, reproducible exploration of the vast state space of concurrent, failure-prone systems, DST empowers engineers to find and fix the rarest and most pernicious bugs before they reach production. Its integration with property-based testing, fuzzing, and fault injection, combined with advances in deterministic hypervisors and simulation frameworks, has made DST accessible to a growing range of systems and organizations. While DST requires significant engineering investment, careful system design, and ongoing maintenance, its benefits in reliability, developer productivity, and user trust are profound. As distributed systems continue to grow in complexity and AI/ML systems become more agentic and autonomous, the need for rigorous, deterministic validation will only intensify. The future of DST lies in deeper integration with formal methods, smarter state-space exploration, and broader applicability to AI/ML and hybrid systems. Organizations that embrace DST, alongside complementary techniques like chaos engineering and formal verification, will be best positioned to deliver robust, trustworthy, and resilient distributed systems in the years ahead. DST Tools and Frameworks Deterministic simulation testing (DST) tooling is still a niche but growing ecosystem. Each has a unique focus — ranging from language-level deterministic runtimes to full-stack hypervisor-based reproducibility. Based on the specific needs a single or combination of them can be picked up. References and Further Reads Taming Chaos — DSTSquashing the Heisenbug with DSTAntithesis — DSTPhil Eaton — DSTRedstone — DST FrameworkResonate — DSTJespen | TickLoom
Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq
Testing POST API requests is an important skill for modern QA and automation engineers working with backend services and microservices. In this article, we’ll explore how to test POST API requests with Playwright and TypeScript, focusing on sending a request body using different approaches. By the end, you will learn how to send POST API requests using the following approaches for adding a request body: JSON Object/Array(data)Stringified JSONJSON FileFaker library Application Under Test We will use the POST /addOrder API of the RESTful e-commerce demo application for this demo. The API schema is provided below: JSON { "user_id": "string", "product_id": "string", "product_name": "string", "product_amount": 0, "qty": 0, "tax_amt": 0, "total_amt": 0 } ] Testing POST API Requests With Playwright TypeScript Playwright provides a powerful request API that allows us to create and manage HTTP request contexts. Let’s walk through how to send POST API requests step-by-step using different approaches for passing the request body: Request Body as JSON Object/Array Let’s send a POST API request with a static JSON Array and verify that the response status code is 201. TypeScript test("POST order details API with static JSON Array", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data:[{ user_id: "1", product_id: "82", product_name: "Cadbury Bar", product_amount: 12, qty: 2, tax_amt: 1, total_amt: 25 }, { user_id: "2", product_id: "80", product_name: "MilkyBar", product_amount: 10, qty: 1, tax_amt: 1, total_amt: 11 } ], }); expect(response.status()).toBe(201); }); Code Walkthrough A Playwright test case for testing the POST API request is created using the request API, which is a built-in Playwright APIRequestContext fixture used to send HTTP requests. Sending the POST Request: The request.post() sends a POST request to the “http://localhost:3004/addOrder/” endpoint. The response received after sending a POST request is stored in the response variable.Passing the Request Body (Static JSON Array): The data property is used to send the request body, which accepts data in JSON Array and JSON object formats. Since the POST /addOrder API accepts an array of order details, allowing multiple orders to be submitted within a single JSON array, the test provides the data in JSON array format.Validating the Response: The response.status() retrieves the HTTP status code. The assertion ensures that the status code 201 is returned, indicating that the orders were successfully created. Request Body Using JSON.Stringify() JSON.stringify() can be used while sending raw payloads, testing malformed JSON, or sending custom-formatted JSON. However, the Content-Type header should be set to application/json so the server correctly interprets the request body as JSON data. Let’s write the same POST /addOrder test using JSON.stringify() and verify that a status code 201 is returned in the response. TypeScript test("POST order details API using JSON.Stringify", async ({ request }) => { const orderData = [{ user_id: "5", product_id: "64", product_name: "Cadbury Mini", product_amount: 5, qty: 3, tax_amt: 1, total_amt: 16 }]; const response = await request.post("http://localhost:3004/addOrder/", { data: JSON.stringify(orderData), headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); The orderData array is defined to contain one order object. Each field represents the order details such as user_id, product_id, product_name, and so on. This is a normal JavaScript object at this stage, not JSON yet. The JSON.stringify(orderData )converts the JavaScript object into a JSON string. As we manually stringify the request payload, we should explicitly set the Content-Type header to application/json. It tells the server to treat the request body as JSON data. Playwright sends the POST API request using the post() method and then validates that the response returns a 201 status code. Request Body as a JSON File Using a JSON file as a request body is a handy approach when testing POST API requests with Playwright TypeScript. It supports large payloads and allows the same file to be easily reused across multiple tests. The following orders.json file will be used as a request payload in the POST /addOrder API: JSON [ { "user_id": "1", "product_id": "79", "product_name": "5 star 10gm Chocobar", "product_amount": 5, "qty": 1, "tax_amt": 0.5, "total_amt": 5.5 }, { "user_id": "2", "product_id": "71", "product_name": "Lindt Milk Chocolate", "product_amount": 15, "qty": 3, "tax_amt": 2.5, "total_amt": 47.5 } ] The following configurations should be in place before we proceed to write the test using a JSON file as the request payload. tsconfig.json JSON { "compilerOptions": { "module": "NodeNext", "moduleResolution": "NodeNext", "resolveJsonModule": true } } We need to ensure that the JSON file is imported first and then attached to the data parameter while sending the POST request. TypeScript import orders from '../test_data/orders.json' with {type: 'json'}; test("POST order details API using JSON file", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data: orders, }); expect(response.status()).toBe(201); }); This test imports test data from an external orders.json file and uses it as the request body for a POST API call in Playwright. The request.post() method sends the imported JSON directly as the payload, and the test verifies that the API returns a 201 status code. Using a JSON file makes it easier to manage, reuse, and update request payloads without modifying the test logic, improving the maintainability and readability of API tests. Request Body With Faker Library Using the Faker library to create a request body for a POST API helps generate realistic and dynamic test data, reduces hardcoded values, and improves test coverage. It is especially useful for simulating real-world scenarios and avoiding duplicate data issues. However, it also has limitations. Since Faker generates random data, it may lead to inconsistent test results if not properly handled, and debugging failures could become more difficult without fixed or reproducible inputs. To use the Faker library, we need to install it first using the following command: Plain Text npm install --save-dev @faker-js/faker Next, let’s create two helper functions: one to create order objects with the required order details, and another to generate an array of orders based on the count provided by the user. TypeScript import { faker } from '@faker-js/faker'; export function createOrderDetails() { const productAmount:number = faker.number.int({ min: 1, max: 100 }); const qty:number = faker.number.int({ min: 1, max: 5 }); const taxAmt:number = faker.number.int({ min: 2, max: 10 }); const totalAmt:number = (productAmount*qty)+taxAmt; return { user_id: faker.number.int({ min: 1, max: 50 }), product_id: faker.number.int({ min: 1, max: 100 }), product_name: faker.commerce.productName(), product_amount: productAmount, qty:qty, tax_amt: taxAmt, total_amt: totalAmt }; } The createOrderDetails() function returns the order details with the required fields. It also calculates the total amount by summing the total product value and tax amount. The values in all fields are updated randomly using the appropriate methods provided by the Faker library. TypeScript import { faker } from '@faker-js/faker'; export function createRandomOrders(count:number) { return faker.helpers.multiple(createOrderDetails, {count}); } The createRandomOrders(count) method accepts a count parameter and returns randomly generated orders that can be directly used in the test. TypeScript return faker.helpers.multiple(createOrderDetails, {count}); This is where the work happens. multiple() is a helper function provided by Faker. Its job is to execute another function multiple times and collect all the results into an array. It takes two arguments: A function to execute repeatedly.An options object specifying how many times to execute it. It asks Faker to call createOrderDetails() exactly count times, collect all the generated orders into an array, and return that array to the caller. TypeScript test('POST order details API using Faker library', async({request}) => { const orderData = createRandomOrders(5); const response = await request.post("http://localhost:3004/addOrder/", { data: orderData, headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); This test sends a POST request to the /addOrder endpoint API using dynamically generated test data from the Faker library. The createRandomOrders(5) function generates an array of 5 random order objects, which are passed as the request body using the data property. The test then verifies that the API responds with a 201 status code, confirming that the orders were successfully created. Summary In this tutorial, we explored multiple approaches to testing POST API requests using Playwright with TypeScript, including sending static JSON payloads, using JSON.stringify, importing data from external JSON files, and generating dynamic test data with the Faker library. We also covered how to handle headers correctly and validate API responses using status code assertions. In my experience, using external JSON files is a practical approach for testing POST API requests because it lets us send bulk data from a single file. Similarly, the Faker library can also be used to generate dynamic test data; however, in some cases, using a third-party library may not be permitted. Ultimately, the software team chooses the approach that best fits their testing strategy. Happy testing!
Unit testing business logic often requires surprisingly little business data. Suppose we want to test a program that loads an order, calculates its total, and rejects it when the amount exceeds a limit. The decision we want to verify is simple: The program loads the specified order.It asks for the order’s total.If the total is too high, it does not approve the order.It returns FALSE. Yet a conventional Java unit test may have to construct an Order. That order may require a customer, line items, currencies, prices, tax information, identifiers, and other objects that have nothing to do with the decision being tested. Builders, fixtures and mocking frameworks reduce the typing, but they do not eliminate the underlying problem: the test must participate in the internal representation of the domain model. BUBAS takes a different approach. Domain objects are opaque to a BUBAS program. Because the program cannot inspect them, a unit test does not need to construct them. It needs only a token. The Business Program Consider this BUBAS program: SQL PROGRAM ApproveOrder(orderId INTEGER, limit DECIMAL) RETURNS BOOLEAN DECLARE purchase Order DECLARE total DECIMAL purchase = LOAD_ORDER(orderId) IF NOT ORDER_WAS_FOUND(purchase) THEN LOG_EVENT "ERROR", "no such order: " + orderId RETURN FALSE END IF total = ORDER_TOTAL(purchase) IF total > limit THEN LOG_EVENT "INFO", "over limit: " + total RETURN FALSE END IF APPROVE purchase RETURN TRUE END. Order is a Java domain type registered by the application embedding BUBAS. The program can store an Order in a variable and pass it to operations that accept an Order, but it cannot access its fields or invoke its methods. There is no expression such as: SQL purchase.customer.account.balance If the program needs information about an order, the application must expose an operation for obtaining it: SQL total = ORDER_TOTAL(purchase) This restriction is primarily an encapsulation mechanism. The business program depends on the vocabulary of its domain rather than on the internal structure of Java objects. It also has an important consequence for testing. Replace the Object With Identity Here is a BUNIT test for the over-limit case: Gherkin PROGRAM OverLimitIsRejected "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "APPROVE _" IS MOCKED ARGUMENT "orderId" IS 42 ARGUMENT "limit" IS 1000.00 RUN RESULT IS FALSE "APPROVE _" WAS NOT CALLED END. The string "o1" is not an order serialized as text. It does not contain an order number, a total or any other property. It is a test token representing one opaque Order. The first mock says: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" When the program calls LOAD_ORDER(42), BUNIT returns the token "o1" in place of the real Java object. The program stores it in purchase. Later it calls: Java ORDER_TOTAL(purchase) The second mock recognizes that same token and returns 1500.00. The program cannot tell that "o1" is not a real Order. It has no operation with which to inspect the object. It can only pass the value back through the vocabulary supplied by the host application. For this test, identity is all the domain object needs. We Are Testing the Conversation A BUBAS business program contains decisions and orchestration. Algorithms, persistence, infrastructure, and domain-object implementations remain in Java. Its unit test should therefore concentrate on questions such as: Which domain operations were invoked?With what arguments?What values did those operations return?Which branch did the program select?Which operations were deliberately not invoked?What result did the program produce? In the example, we do not test how ORDER_TOTAL calculates a total. That belongs in the Java test for the implementation of ORDER_TOTAL. We test what the business program does when ORDER_TOTAL reports 1500.00. This division gives us two focused tests rather than one oversized test: Java tests verify the individual domain operations.BUNIT tests verify how a business program coordinates them. The BUNIT test documents the business scenario directly. An order identified by 42 exists, its total is 1500.00, the approval limit is 1000.00, and the program must not approve it. The test does not explain how to manufacture an object graph that produces those facts. Opacity Buys Mockability Mocking domain objects in a general-purpose language is often difficult precisely because the production code can observe so much about them. It may call methods, inspect nested objects, compare values, serialize the object or pass it to code that expects a particular implementation. A substitute must reproduce every observable property used along the tested path. An opaque BUBAS value has only the observations provided by the registered vocabulary. If the vocabulary exposes ORDER_TOTAL, then the mock controls the answer to ORDER_TOTAL. If it does not expose the customer’s internal account object, neither the program nor the test needs to know that such an object exists. The object boundary and the testing boundary are the same boundary. This is stronger than merely saying that business programs should avoid inspecting domain objects. They cannot inspect them unless the embedder deliberately provides an operation that does so. Consequently, a token can stand in for any opaque value as long as the mocks define how the exposed operations respond to it. Multiple objects require only multiple identities: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "LOAD_ORDER" WITH ARGS(43) RETURNS "o2" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "ORDER_TOTAL" WITH ARGS("o2") RETURNS 200.00 The test describes the distinctions that matter without constructing either order. The Test Uses the Real Language A dangerous form of mocking creates a second, simplified interface used only by tests. Eventually, the production vocabulary changes while the test vocabulary does not. BUNIT does not compile the business program against a parallel language. The program under test is compiled against the real sealed BUBAS language. Mocking happens later, at dispatch. Therefore, the test cannot silently keep using an operation that no longer exists in the production language. Nor can it casually return a value of the wrong BUBAS type. Before executing a test, BUNIT checks the mocks and the test configuration. It can report problems such as: A mock declared with the wrong number of arguments;A mock returning a value incompatible with the real operation;An argument supplied for a parameter the program does not accept;A mocked command that should initialize a variable but does not provide its value. The test reports these errors before the business program runs. The test remains artificial — as every unit test is — but it is artificial inside the actual language contract. Do Not Assert Everything A test becomes fragile when it records every interaction, whether or not that interaction matters to the scenario. BUNIT allows an expectation to specify only the relevant part of a call. For example: Plain Text "LOG_EVENT _, _" WAS CALLED WITH ARGS("INFO", CONTAINS("over limit")) The test requires an informational log message containing "over limit". It does not require the complete message to remain byte-for-byte identical. Similarly: Plain Text "APPROVE _" WAS NOT CALLED expresses the important negative requirement without inventing an Order merely to compare it with another Order. The purpose is not to reproduce the execution trace. It is to state the observable facts that define the business case. What This Does Not Test Opaque tokens do not prove that the Java implementation of LOAD_ORDER returns the right order. They do not prove that ORDER_TOTAL calculates taxes correctly or that APPROVE commits a transaction. Those operations require their own Java unit and integration tests. BUNIT tests the program at the orchestration boundary. This makes it possible to test business decisions without databases, service containers, or complete domain-object graphs, but it does not replace testing below or beyond that boundary. Nor does BUNIT make every Java application automatically testable. The application developer first has to expose a suitably designed vocabulary. If one enormous operation performs loading, calculation, approval and notification internally, BUNIT can mock that operation but cannot test the decisions hidden inside it. Testability therefore provides feedback about vocabulary design. Operations should represent meaningful domain capabilities at the level where business programs genuinely make choices. The Deeper Result Opaque domain types may initially look like a limitation. The program cannot examine its own values freely. It has to ask the vocabulary to interpret them. That limitation creates a clean separation: Java owns domain representation and implementation.BUBAS owns orchestration and decisions.BUNIT replaces domain capabilities at that same boundary.Tokens replace complex objects with identity when identity is all the test requires. The production program becomes independent of domain-object structure. The unit test inherits that independence. We do not need a fake Order with a fake customer containing fake line items whose prices happen to add up to 1500.00. For this business decision, we need only to say: Plain Text "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 The business program never needed to know what was inside the order. Neither does its test. The detailed code and the BUBAS framework are available as open source at https://github.com/verhas/bubas.
One of the most important parts of API test automation is validating the response body to ensure data integrity. This step plays a key role in functional API testing, as it helps confirm that the API is returning the right data in the expected format. Response body validation isn’t limited to a specific request type; it applies equally to POST, GET, PUT, and PATCH APIs. The same validation approach can be used for any API response to verify the data returned by the service. Playwright offers multiple ways to validate response bodies. In this tutorial, I’ll walk you through these approaches to help you efficiently perform assertions on the response data using best practices. Checkout the previous tutorial blog to learn about Installation, the demo application, and how to send GET API requests with Playwright. How to Verify the Response Structure Response structure checks ensure that an API consistently returns data in the expected format, protecting the contract between backend services and their consumers. They help catch breaking changes early, such as missing or renamed fields, even when the API still returns a successful status code. TypeScript test("GET Order details and perform structure check", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toHaveProperty("message"); expect(responseBody).toHaveProperty("orders"); expect(responseBody.orders[0]).toHaveProperty("id"); expect(responseBody.orders[0]).toHaveProperty("product_name"); }); This test focuses on validating the structure of the API response. It validates that the response body contains the expected top-level keys and that each order object includes the required fields. Basic Assertions The basic assertions validate API success and data presence, making them a good first layer of verification before deeper structure or data-level checks. TypeScript test("Get order details and perform basic level verification", async ({ request, }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(Array.isArray(responseBody.orders)).toBeTruthy(); expect(responseBody.orders.length).toBeGreaterThan(0); }); This test performs a basic level check to confirm that the endpoint works as expected and returns the expected data in the response. After parsing the response body, the assertions focus on the following essential basic-level checks: TypeScript expect(responseBody.message).toBe("Order found!!"); The above line of code verifies that the API returns the expected message text in the response body. TypeScript expect(Array.isArray(responseBody.orders)).toBeTruthy(); This line of code ensures that the orders field in the response is an array, validating the basic response format. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This part of the test confirms that at least one order is returned in the orders array, ensuring the response contains required data. How to Verify Response Data With Details Validating the actual data returned in the response is essential to ensure that the API response contains the correct values. TypeScript test("Get order and verify order details", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); }); The following code ensures that the response has a valid identifier and it is not missing or empty. TypeScript expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); This check is required because the API generates the order ID when a new order is created in the system. It ensures that the “id” field has a valid value generated and assigned to it, since this “id” is used to retrieve, update, or delete order data. TypeScript expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); These statements assert that the order details are retrieved correctly for the respective request. The “user_id” - “1” was sent in the request, and verifying it in the response, along with the other order details such as “product_id” and “product_name,” ensures that the correct data is returned. How to Verify Response Data by Matching Objects and Arrays Playwright allows response data verification by matching objects and arrays partially within the API response. This approach is useful because it makes tests more flexible and confirms that the API returns the correct data structure and values. TypeScript test("Get order and verify matching object and array", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toMatchObject({ message: "Order found!!", orders: expect.arrayContaining([ expect.objectContaining({ product_id: "79", product_name: "5 star 10gm Chocobar", product_amount: 5, qty: 1, tax_amt: 0.5, total_amt: 5.5, }), ]), }); }); In this test, the toMatchObject assertion verifies that the response contains a “message” with the expected value “Order found!!” and an orders array. Within the array, "expect.arrayContaining" ensures that at least one order matches the expected data, while "expect.objectContaining" verifies only the values in the specified fields of that order. Using Best Practices to Perform Assertions Best practices create stable, maintainable API automation tests by combining basic checks with flexible data matching. TypeScript test("Get Order details API test with best practice", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); }); The test sends a GET request to fetch order details for “user_id”-“1". The use of failOnStatusCode: true ensures the test fails immediately if the API does not return a 2xx status code. The response is then parsed into a JSON object for validation. The assertions are structured in layers: TypeScript expect(responseBody.message).toBe("Order found!!"); This assertion verifies the message text, confirming that the API returns the correct message when an order is found. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This statement ensures meaningful data is returned and avoids false positives when the array is empty. TypeScript expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); The final part of the code performs the final assertion using arrayContaining and objectContaining to verify that at least one order has the expected “id” and “product_name”, without asserting every field. These layered validations improve clarity by verifying structure, data presence, and key data values in sequence. Extracting Data From the Response Extracting data from the API response is a common and widely used pattern in API test automation. It is important in multiple ways, such as reusing the data in further tests for dynamic testing and end-to-end validation. TypeScript test('Get order details and extract the order id', async({request}) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); const product_name = order.product_name console.log(product_name) }); This test sends a GET API request and performs basic validations to ensure the API response is reliable. TypeScript const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); The code above extracts the “order_id” from the order object in the response. Before accessing it, an assertion is made to verify that the value is not null. Finally, the value of the order_id is printed in the console. TypeScript const product_name = order.product_name console.log(product_name) Similarly, other values, such as product_name, can also be extracted. Attaching the Response Body to the Playwright Report The Playwright report, by default, shows the steps executed, the number of tests run, pass/fail status, and time taken to run the tests. However, it does not attach the response body to the test report. Attaching the response body to the report improves visibility and makes the test report more informative and transparent. The following code shows how to extract the required metadata and attach it to the Playwright report. TypeScript test("Get order details API and attach the response details to the report", async ({ request, }, testInfo) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, }); expect(response.status()).toBe(200); const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); const fullResponse = { status, statusText, headers, body, }; await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); }); The testInfo is a built-in Playwright fixture and provides utilities to manage and inspect test execution, such as attaching files to reports, updating test timeouts, and identifying the currently running test. The following lines of code extract the response metadata, such as the status code, status text, headers, and response body. TypeScript const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); Next, let’s combine all response details and create a single object containing: Status codeStatus textHeadersResponse body TypeScript const fullResponse = { status, statusText, headers, body, }; Finally, let’s attach these details to the report using the testInfo.attach() method as shown below: TypeScript await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); The testInfo.attach() adds an attachment to the Playwright report. The attach() method has 3 parameters: Name of the attachment: The first parameter is the name, “Full API Response”, that will be shown for the attachment.Body of the attachment: The second parameter is for the body of the attachment. The JSON.stringify(fullResponse, null, 2) has 3 arguments. The first argument converts the fullResponse object into a readable, pretty-formatted JSON. The second argument is the replacer, which is null. It ensures that all properties from the fullResponse object are included as they are, without modifying anything. The third argument controls pretty-printing. Here, “2” means indent nested JSON by 2 spaces.Content type: This parameter ensures that the report treats the attachment as JSON. The following screenshot is generated after the tests are run: Test Execution Running the tests in Playwright is simple and easy. We can run the following command from the terminal: Plain Text npx playwright test To generate the report, the following command can be used: Plain Text npx playwright show-report Summary Playwright provides multiple approaches, including structure checks and matching objects and arrays for verifying response data. The right strategy should be chosen based on your project’s requirements. Based on my experience, combining response structure checks with response data validation, including the matching object and array strategy, can be used as an effective approach for validating API responses. Happy testing!
Production failures often contain enough evidence to explain what went wrong, but not enough structure to become an executable test. A trace may expose the failing request path, a log may contain the exception, and downstream spans may reveal the dependency response that triggered the defect. The useful engineering step is to transform that evidence into a deterministic regression test rather than another incident summary. Recent bug-reproduction systems follow the same principle that a useful reproducer should fail on the buggy revision for the reported reason and become passing evidence after the defect is fixed. Issue2Test and ReProAgent both use execution feedback instead of treating test generation as a single prompt-and-response operation. Start From the Incident Evidence The agent should begin from a machine-readable incident envelope, not a copied stack trace. OpenTelemetry’s stable log data model includes TraceId and SpanId, while its exception conventions associate exception records with the corresponding span context. W3C Trace Context standardizes traceparent for propagating trace identity across service boundaries. Those identifiers allow the failing execution path to be reconstructed without forwarding an entire observability dataset to a model. A small adapter can convert an alert into the minimum evidence required by the agent: Java FailureContext buildContext(Incident incident) { Trace trace = telemetry.getTrace(incident.traceId()); Span failed = trace.failedSpan(); return new FailureContext( failed.operation(), failed.exception(), trace.parentPath(failed), trace.downstreamCalls(failed), repository.revision(incident.deploymentId())); } The deployed revision is essential. A regression test generated against current source can target code that has already moved away from the production state. The incident should therefore resolve to the commit, image digest, or equivalent immutable revision that produced the telemetry. The trace supplies runtime evidence, and the repository supplies the code that interpreted it. Telemetry also requires reduction before model access. Request bodies, authorization headers, customer identifiers, and database values are rarely necessary to reproduce control flow. OpenTelemetry documents Collector processors to remove attributes, filter records, redact attributes, and transform values before export. Those controls should run before failure context reaches the agent rather than relying on a model to ignore sensitive fields. Reduce the Failure to Executable Context Raw traces are too broad for test generation. The agent needs a compact slice containing failing application frames, the request shape, relevant downstream interactions, and nearby tests that define local conventions. ReProAgent’s 2026 design separates bug localization, root-cause analysis, test planning, and test generation, combining repository retrieval with runtime interaction. Its results support treating reproduction as a staged, tool-using process rather than direct code completion. For a checkout failure, an error span may show InventoryClient.reserve() followed by a NullPointerException after the inventory service returned HTTP 503. Retrieval should locate InventoryClient, the calling checkout path, exception mapping, and existing checkout tests. Unrelated controllers, persistence code, and complete trace payloads add noise without strengthening the reproducer. The resulting agent input can be expressed as an explicit contract: Java TestRequest request = new TestRequest( context.failureFingerprint(), context.relevantSource(), context.relatedTests(), context.downstreamResponses(), "Generate one deterministic JUnit regression test. " + "Do not modify production code. Do not assert the observed bug as correct behavior." ); That final constraint is critical. A model can produce a test that asserts NullPointerException simply because production emitted it. Such a test would pass on the buggy implementation and preserve the defect. Bug-reproduction benchmarks instead use fail-to-pass behavior where the test fails on the pre-fix revision and passes after the correcting patch. Recent research on LLM repair validation also finds that passing executions can provide little bug-discriminating evidence, making differential validation important. Generate the Test Against the Intended Contract The oracle should come from repository evidence rather than model invention. Existing tests, API specifications, exception policies, sibling implementations, and documented response contracts can establish intended behavior. When those sources conflict, the candidate should remain unresolved instead of receiving a fabricated assertion. Consider a production failure where inventory returned 503 and checkout converted a missing response body into an internal NullPointerException. Existing endpoint tests may establish that unavailable dependencies map to a stable 503 response with an INVENTORY_UNAVAILABLE code. The generated regression test can encode that contract while reproducing the recorded dependency behavior: Java stubFor(post(urlEqualTo("/inventory/reservations")) .willReturn(aResponse() .withStatus(503) .withBody("{\"code\":\"overloaded\"}"))); mockMvc.perform(post("/orders") .contentType("application/json") .content(failureRequest)) .andExpect(status().isServiceUnavailable()) .andExpect(jsonPath("$.code").value("INVENTORY_UNAVAILABLE")); WireMock can match HTTP requests and return predefined responses, and it supports fixed or randomized delays and lower-level fault simulation. That allows a recorded external condition to become a deterministic test setup rather than a dependency on a live production service. Close the Loop With Execution Feedback Generation should be treated as the first candidate, not the final artifact. Issue2Test refines tests using compilation and runtime feedback, while ReProAgent includes runtime interaction throughout reproduction. A practical agent should compile and execute every candidate in an isolated checkout of the incident revision. Java TestCandidate refine(TestCandidate candidate, FailureContext context) { for (int attempt = 0; attempt < 4; attempt++) { TestRun run = sandbox.run(context.revision(), candidate); if (run.compiles() && reproduces(run, context)) return candidate; candidate = model.revise(candidate, run.diagnostics(), context); } return TestCandidate.rejected(); } The reproduces check should be stricter than “test failed.” It can verify that the expected application path was reached, the recorded downstream condition was exercised, and the observed exception or response fingerprint overlaps the incident. Compilation failures feed back into correction, a test that fails before reaching the target path is rejected and a test that passes on the buggy revision is not a reproducer. Once a fix exists, the same test should run against both revisions. ReProAgent defines fail-to-pass rate around exactly this distinction: failure on the buggy state and success after the issue-resolving patch. Differential execution is stronger evidence than asking a model whether generated code appears correct. Make the Test the Durable Artifact After deterministic replay, the reproducer can enter the normal test suite. JUnit treats failed assertions and uncaught exceptions as test failures, so ordinary CI can enforce the regression once the test is valid. Normal execution should require neither production telemetry nor another model call, and incident secrets should never be embedded in the generated fixture. A practical CI handoff can also preserve provenance without preserving raw incident data. A small metadata record can contain the incident identifier, source revision, generated test path, reproduction fingerprint, and validation command. That record makes regeneration and review easier while keeping the committed test independent of the observability backend. The test itself remains the executable source of truth. In practice, the generated test is verified under strict CI controls before ever reaching the main suite. The agent’s changes (adding the new test) occur on an isolated branch or worktree, and the CI pipeline runs git diff to confirm that only test files were created or modified, any application code changes cause an immediate failure. The test is then run against the original codebase to confirm it reproduces the production failure, and again against the patched build to ensure it now passes. Any anomaly (for example, the test accidentally passing on the buggy code or still failing after the fix) triggers a manual review. Meanwhile, any necessary fixtures from the incident (such as specific database records or request parameters) are set up in the test so it precisely mirrors the failure scenario. Metadata from the failure (stack trace, error message, etc.) is included in the commit or PR for traceability. This enforces that each generated test is precise and verifiable in CI before the developer ever sees it. Production observability becomes substantially more valuable when failures can be converted into executable evidence. The reliable pattern is to correlate telemetry to the deployed revision, reduce that evidence to the failing path, derive assertions from existing contracts, generate a deterministic test, and repeatedly execute it until the production failure is faithfully reproduced. The final acceptance criterion is demanding but clear: the test must fail for the real bug, pass after the real fix, and remain safe enough to run on every future change. That turns an AI debugging agent from a code generator into a controlled mechanism for converting operational failures into permanent regression protection.
When a QA team is asked, "How confident are we in this AI system?" the instinctive answer is to write more test cases. If 100 test cases gave us some confidence, surely 1,000 will give us ten times more. This instinct is deeply wired into traditional software testing, where every additional test case can, in principle, catch a bug the others missed. For AI systems, this instinct is not just inefficient — it is often mathematically wrong. Adding more test cases the wrong way can leave you with less real confidence than a much smaller, carefully sampled set. Understanding why requires looking at what a test case is actually measuring when the system under test is probabilistic rather than deterministic. This article walks through the statistical reasoning that should drive AI test suite design, and shows why smart sampling — not raw volume — is what actually buys you confidence in an AI system's behavior. What a Test Case Measures for Deterministic Software For traditional software, a test case answers a yes-or-no question: given this input, does the system produce the expected output? Each test case is independent evidence about a specific code path. Adding a new test case that exercises a previously untested branch genuinely adds new information, because the system's behavior on that branch was, until that test ran, completely unknown. This is why test coverage metrics — line coverage, branch coverage — are meaningful for deterministic software. Coverage tells you what fraction of the system's possible behavior has been directly observed at least once. Why the Same Logic Breaks for AI Systems An AI system does not have a fixed, enumerable set of behaviors the way a codebase has a fixed set of branches. Its behavior is a distribution — a probability of producing each possible output for a given input, and often a probability that itself shifts slightly across runs, contexts, and time. When you write a single test case for an AI system — one prompt, one expected answer — you are not observing "the" behavior of the system for that input. You are observing one sample drawn from a distribution of possible behaviors. Running that exact same test case again might draw a different sample from the same distribution. This changes the statistical meaning of a test case entirely. Here's where everything changes: a single AI test case is no longer a definitive answer — it is just one observation from a much larger behavioral distribution. In deterministic testing, one test case answers one question conclusively. In AI testing, one test case gives you one data point toward estimating a probability — and a single data point tells you almost nothing about a probability distribution. The Confidence Interval Problem Consider a concrete example. You want to know: does this AI-powered customer support assistant give a policy-compliant answer to refund questions at least 95% of the time? You write 20 refund-related test cases, run them, and 19 pass. That is a 95% pass rate — right at your target. Statistically, this result is far weaker evidence than it feels. With only 20 samples, the 95% confidence interval around your observed 95% pass rate is wide — the true underlying success rate could plausibly be anywhere from roughly 75% to 100%. Twenty test cases have not told you the system meets your bar. They have told you the system's true performance is consistent with a wide range that happens to include your bar. To narrow that confidence interval meaningfully — to actually distinguish a system that is truly 95% reliable from one that is truly 85% reliable — requires a sample size in the hundreds, not tens, for that specific scenario category. In practice, engineering teams commonly use 95% confidence when estimating AI system reliability, although higher confidence levels may be appropriate for safety-critical applications. This is the first place where "more test cases" and "smart sampling" start to diverge: you need enough samples per scenario category to say anything statistically meaningful about that category, and no amount of test cases in unrelated categories substitutes for that. There's a simple piece of math behind this that's worth internalizing, because it explains why the problem doesn't go away just by writing more tests. The width of a confidence interval shrinks in proportion to the square root of your sample size, not in direct proportion to it: Plain Text Confidence Interval width ∝ 1 / √n Doubling your sample size does not double your confidence — it only shrinks your uncertainty by about 30%. Quadrupling it cuts uncertainty in half. This is the quiet reason volume-based testing feels productive while delivering diminishing returns: the tenth test case in a category buys you real statistical ground, but the two-hundredth buys you very little compared to what it costs to write and run. What This Looks Like in Practice The contrast between the two approaches is easiest to see side by side. Plain Text TRADITIONAL TESTING 1,000 test cases spread evenly │ ▼ "Coverage" (looks thorough, says little about any one scenario) SMART SAMPLING Scenario A (high risk) → 120 samples Scenario B (high risk) → 80 samples Scenario C (medium risk) → 60 samples Scenario D (low risk) → 20 samples │ ▼ Confidence (narrow intervals where it matters most) The traditional approach optimizes for a number that looks reassuring on a dashboard. The sampling approach optimizes for statistical confidence where business risk is highest — and is honest about where confidence is intentionally looser. Where Volume Actually Hurts Here is the counterintuitive part. Teams often respond to this problem by writing more test cases — but they add them across many different scenario categories rather than deepening any single one. The result is a suite with 500 test cases, twenty scenario categories, and roughly 25 samples per category — still not enough to draw a confident conclusion about any individual category, while creating the appearance of a large, thorough suite. This is worse than it sounds, because a large test suite carries real costs. It takes longer to run, which slows down CI/CD feedback loops. It takes longer to maintain, since ground-truth answers for AI systems need periodic review as policies and knowledge bases change. And critically, it creates false confidence — a dashboard showing "500 tests, 98% pass rate" reads as strong evidence to a stakeholder, when the underlying statistics may not support that read at all for any specific scenario that stakeholder actually cares about. What Smart Sampling Looks Like Instead Smart sampling starts from a different question: not "how many test cases can we write," but "what decision do we need statistical confidence about, and how many samples does that decision actually require." The first step is defining scenario categories that map to real business risk — refund policy questions, account security questions, product availability questions — rather than categories that map to convenient technical groupings like "single-turn queries" versus "multi-turn queries." Risk-aligned categories are what stakeholders actually need confidence about. The second step is stratified sampling within each category: generating semantically varied inputs that probe the same underlying scenario from different angles — different phrasings, different levels of ambiguity, different amounts of context — rather than many near-duplicate test cases that differ only in superficial wording. Ten semantically diverse samples of a scenario carry more statistical information than fifty near-identical restatements of the same question, because the near-identical restatements are highly correlated with each other and do not independently sample the underlying distribution. The third step is allocating sample size deliberately by risk. A scenario category with high business consequence — anything touching financial transactions, medical guidance, or legal disclosures — warrants a large enough sample to produce a narrow confidence interval, potentially hundreds of cases. A low-consequence category, such as a cosmetic formatting preference, can be validated adequately with a much smaller sample. Treating every category with the same sample size wastes effort on low-risk scenarios while under-sampling high-risk ones. The fourth step is repeated sampling over time rather than only at initial test design. Because AI system behavior can drift, a sample that gave a narrow, confident interval six months ago does not guarantee the same interval holds today. Smart sampling treats the sample size and scenario allocation as something to periodically re-justify against current production data, not a decision made once and left untouched. The same sampling principles apply directly to retrieval-augmented generation (RAG) systems, which now sit behind most enterprise AI assistants. In a RAG pipeline, response correctness depends on both the generative model and the quality of what gets retrieved, so a scenario category isn't fully sampled unless it captures variation in retrieval outcomes too — cases where the right document is retrieved, cases where a close-but-wrong document is retrieved, and cases where retrieval comes back empty. Treating "RAG testing" as one category instead of a set of retrieval-quality-weighted sub-scenarios is one of the most common places teams under-sample without realizing it. The same statistical reasoning also applies to autonomous AI agents. Because agents make sequential decisions, validation must consider the probability distribution across complete workflows rather than evaluating each individual step in isolation. A Practical Illustration Suppose an enterprise AI validation team has a fixed budget of 300 test executions per CI/CD run — a real constraint, since each execution costs inference time and, for hosted models, direct API cost. A volume-first approach might spread these 300 across 30 scenario categories, roughly 10 samples each — statistically too thin to draw confident conclusions about any single category. A risk-based sampling approach might instead allocate 80 samples to the three categories touching financial and account-security actions, 40 samples each to five categories with moderate business consequence, and 10 samples each to the remaining ten low-consequence categories. The total sample budget is unchanged at 300, but the confidence intervals for the categories that actually matter to the business are now meaningfully narrower, while low-risk categories still receive baseline coverage rather than none at all. This is the essence of smart sampling: the same testing budget, reallocated according to statistical need and business risk, rather than spread evenly across categories regardless of consequence. The Broader Principle The deeper lesson here extends beyond test case counting. AI validation, as a discipline, has to import statistical thinking that traditional software testing rarely required, because traditional testing dealt with deterministic systems where a single well-chosen test case could conclusively answer a question. AI systems require thinking in terms of distributions, confidence intervals, and sample sizes — the vocabulary of applied statistics rather than the vocabulary of test coverage. Teams that continue to measure AI test suite quality purely by test case count will keep producing dashboards that look reassuring and mean less than they appear to. Teams that shift to measuring statistical confidence per risk-weighted scenario category will produce smaller, faster, and — despite being smaller — genuinely more informative test suites. AI systems are not validated by counting test cases. They are validated by measuring uncertainty. The future of AI quality engineering belongs to teams that measure confidence — not coverage. As enterprise AI systems become increasingly autonomous, statistical validation will become as fundamental to software quality engineering as code coverage is today.
Sr. QA Lead Manager,
3Pillar
Lead Engineer,
Technical University of Crete
Blogger, QA, Mentor, Trainer,
Freelancer