DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Data Engineering

Welcome to the Data Engineering category of DZone, where you will find all the information you need for AI/ML, big data, data, databases, and IoT. As you determine the first steps for new systems or reevaluate existing ones, you're going to require tools and resources to gather, store, and analyze data. The Zones within our Data Engineering category contain resources that will help you expertly navigate through the SDLC Analysis stage.

Functions of Data Engineering

AI/ML

AI/ML

Artificial intelligence (AI) and machine learning (ML) are two fields that work together to create computer systems capable of perception, recognition, decision-making, and translation. Separately, AI is the ability for a computer system to mimic human intelligence through math and logic, and ML builds off AI by developing methods that "learn" through experience and do not require instruction. In the AI/ML Zone, you'll find resources ranging from tutorials to use cases that will help you navigate this rapidly growing field.

Big Data

Big Data

Big data comprises datasets that are massive, varied, complex, and can't be handled traditionally. Big data can include both structured and unstructured data, and it is often stored in data lakes or data warehouses. As organizations grow, big data becomes increasingly more crucial for gathering business insights and analytics. The Big Data Zone contains the resources you need for understanding data storage, data modeling, ELT, ETL, and more.

Data

Data

Data is at the core of software development. Think of it as information stored in anything from text documents and images to entire software programs, and these bits of information need to be processed, read, analyzed, stored, and transported throughout systems. In this Zone, you'll find resources covering the tools and strategies you need to handle data properly.

Databases

Databases

A database is a collection of structured data that is stored in a computer system, and it can be hosted on-premises or in the cloud. As databases are designed to enable easy access to data, our resources are compiled here for smooth browsing of everything you need to know from database management systems to database languages.

IoT

IoT

IoT, or the Internet of Things, is a technological field that makes it possible for users to connect devices and systems and exchange data over the internet. Through DZone's IoT resources, you'll learn about smart devices, sensors, networks, edge computing, and many other technologies — including those that are now part of the average person's daily life.

Latest Premium Content
Trend Report
Cognitive Databases, Intelligent Data
Cognitive Databases, Intelligent Data
Trend Report
Platform Engineering and DevOps
Platform Engineering and DevOps
Refcard #291
Code Review Core Practices
Code Review Core Practices
Refcard #403
Shipping Production-Grade AI Agents
Shipping Production-Grade AI Agents

DZone's Featured Data Engineering Resources

Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code

Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code

By Uthej Mopathi DZone Core CORE
AI-generated software needs provenance that survives beyond the chat window. A code review can show what changed, but it rarely shows which model produced a fragment, what prompt and repository context influenced it, which agent or tool executed the change, which intent was being implemented, or whether the recorded history was altered later. Provenance fills that gap by treating generation as a supply-chain event rather than an ephemeral interaction. The core idea is already established in adjacent standards where W3C PROV models provenance through entities, activities, and agents, while SLSA records how software artifacts were produced so downstream consumers can verify expected processes and inputs. Generation Provenance as Engineering Metadata The first design rule is to separate authorship from provenance. Provenance answers where code came from and how it was produced, and it does not, by itself, determine legal ownership. The U.S. Copyright Office states that generative-AI output is copyrightable only where sufficient human-authored expressive elements exist, and that prompting alone is not enough. Employment agreements, contributor agreements, licenses, and jurisdictional law still govern ownership questions. Provenance instead supplies evidence for attribution, review, audit, and accountability. A minimal record should bind the generated artifact to the model provider, model identifier and revision, agent identity and version, prompt digest, context digests, execution trace, repository commit, intent digest, timestamp, and approving human or service identity. Hosted model aliases can change over time, so a provider-returned model identifier or immutable deployment revision is preferable to a friendly model name alone. The record should also contain cryptographic digests for generated files so later edits cannot silently inherit stale provenance. JSON { "artifact": "src/billing/CancelService.java", "sha256": "7e91...c42a", "commit": "9f3c1ad", "model": {"provider": "acme-ai", "id": "code-model", "revision": "2026-08-14"}, "agent": {"id": "repo-agent", "version": "3.7.2"}, "prompt": "sha256:18ab...90ef", "context": ["git:9f3c1ad^", "sha256:44c2...bb10"], "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1...", "approvedBy": "team:payments-reviewers" } Git commit trailers provide a low-friction place to attach pointers because Git supports structured token-value trailers at the end of commit messages. The commit should store references and digests rather than sensitive prompts themselves. Plain Text AI-Provenance: sha256:5df0...a992 AI-Model: acme-ai/code-model@2026-08-14 AI-Agent: [email protected] Prompt-Digest: sha256:18ab...90ef Context-Digest: sha256:44c2...bb10 AI-Trace: urn:uuid:2cf1... Intent-Digest: sha256:a771...0d61 File-level attribution can use a compact pointer rather than duplicating the full record. A generated region can carry a comment such as // ai-provenance: urn:gen:2cf1..., while the referenced sidecar record maps that generation to file hashes and, when needed, line ranges. This keeps source readable and prevents model metadata from becoming scattered, inconsistent comments. From SBOM to Generation BOM AI-generated code needs an additional description of the generation process. CycloneDX already supports source code, machine-learning models, component provenance, formulation describing how objects were created, and citations that attribute supplied information to entities or processes. Its ML-BOM capability records models, datasets, configurations, and provenance, while the earlier model-card framework established the broader practice of documenting model identity, intended use, evaluation, and limitations. A generation BOM can therefore be implemented as a small signed sidecar linked to the repository and release artifact, rather than inventing a second source-control system. JSON { "bomFormat": "GenerationBOM", "specVersion": "0.1", "subject": {"path": "src/billing/CancelService.java", "sha256": "7e91...c42a"}, "generator": {"agent": "[email protected]", "model": "acme-ai/code-model@2026-08-14"}, "inputs": {"prompt": "sha256:18ab...90ef", "context": ["sha256:44c2...bb10"]}, "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1..." } The intent digest is especially important. A prompt records instructions presented to a model, but an intent contract records the behavior that must remain true after generation. Such a contract can contain permitted change scope, protected behaviors, security constraints, and acceptance criteria. Provenance then connects the produced code not merely to an AI request, but to a reviewable engineering objective. This mirrors data-lineage systems such as OpenLineage, which associate runs, jobs, datasets, and extensible facets so downstream analysis can reconstruct how an output was produced. Artifact hashes alone cannot establish reproducibility when generation depends on mutable infrastructure. Provenance should therefore bind execution parameters such as model configuration, decoding settings, tool versions, retrieval indexes, and policy revisions. Capturing these values converts provenance from a historical label into a verifiable reconstruction boundary for later audits and incident analysis. Tamper Evidence and CI Enforcement Metadata becomes trustworthy only when alteration is detectable. SLSA explicitly treats provenance authenticity and digital-signature verification as mechanisms for detecting tampering, and recommends approaches that improve compromise detection, including transparency logs. Sigstore provides signing with short-lived identity-bound certificates and records signing events in Rekor, an append-only transparency log. A provenance file can be signed as a blob during CI: Shell cosign sign-blob \ --bundle generation-provenance.sigstore.json \ generation-provenance.json Verification should occur before merge or release, not after an incident. SLSA similarly emphasizes that provenance has little value unless a consumer verifies it against expected properties. Shell provctl verify \ --commit "$GIT_COMMIT" \ --require-model \ --require-context-digest \ --require-intent \ --require-signature \ --max-unattributed-lines 0 A practical gate should reject AI-marked changes when the artifact digest no longer matches, required model or agent fields are absent, the provenance signature fails, the intent contract is missing, or the trace cannot be resolved. Human-edited code should not be forced into artificial AI attribution; instead, the policy should distinguish generated, transformed, and manually authored regions. Execution traces can preserve tool calls, retrievals, test runs, and agent steps and are designed to make software-supply-chain steps transparent by recording what happened, by whom, and in what order, and their runtime-trace predicate can describe system events associated with a supply-chain step. Runtime verification closes another gap. Provenance can prove which generation path produced a deployment, but not that the resulting behavior remains correct under production conditions. Release telemetry should therefore link runtime incidents back to commit, provenance record, model revision, and intent contract. That correlation turns an AI-related defect from an unstructured forensic exercise into a query over lineage. Accountability Without Capturing Everything Capturing every prompt verbatim is usually the wrong default. Prompts and retrieved context may contain credentials, personal data, proprietary code, customer information, or licensed material. A safer design stores encrypted source material in an access-controlled evidence store and places digests, object references, retention class, and classification labels in Git-visible provenance. High-sensitivity environments can retain only keyed digests and approved summaries where reproduction is less important than proof of correspondence. Storage and performance costs also require boundaries. Full agent traces can be large, while line-level metadata can become noisy after refactoring. The durable unit should normally be a generation event bound to artifact digests and commits, with finer-grained ranges reserved for high-risk code. Developer ergonomics matter equally, as provenance capture should be automatic in IDE agents, repository bots, and CI runners rather than dependent on manual form filling. Regulation strengthens the case for disciplined records without creating a universal rule that every AI-generated source line must carry a label. The EU AI Act requires general-purpose AI model providers to maintain technical documentation and copyright-compliance policies, while NIST SP 800-218A extends secure development practices specifically for generative AI across the software lifecycle. These frameworks reinforce documentation, traceability, and governance, but a code-provenance system should be treated as engineering evidence rather than a substitute for legal analysis. AI-generated code should enter a repository with the same expectations applied to any other supply-chain artifact: origin, inputs, process, identity, integrity, and approval must be recoverable later. The strongest implementation is not a comment saying that AI was used, but a signed provenance chain linking model and agent identity, prompt and context digests, execution trace, intent contract, commit, generated artifact, review decision, and runtime evidence. Teams adopting AI-assisted development should make that chain automatic, verify it in CI, protect sensitive evidence separately, and fail closed for unattributed high-risk changes. That converts provenance from documentation into an enforceable engineering control and makes accountability possible long after the generation session has disappeared. More
Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

By Ai Cerrudo
Meta has spent years helping businesses advertise on Facebook and Instagram and talk to customers through WhatsApp. Now it wants its AI working more deeply within those businesses. The company launched Meta Enterprise Platform on Monday, a new business focused on turning its AI models, agents, coding tools, and infrastructure into products companies can deploy themselves. CEO Mark Zuckerberg called the platform the "next major pillar" of Meta's business. Its initial technology stack includes the Muse agent, Meta Business Agent, Muse API, and Muse Code. The move puts Meta more directly into an enterprise AI market already crowded with Microsoft, Salesforce, Google, ServiceNow, and other vendors racing to make AI agents part of everyday business operations. Meta is turning its AI stack into a business platform Meta says it already helps hundreds of millions of businesses reach customers. Enterprise Platform is an attempt to expand that relationship beyond advertising and messaging. Rather than introducing a single standalone enterprise application, Meta plans to bring together several components of its AI portfolio. Muse is Meta's general-purpose AI agent, capable of taking actions such as sending emails, booking travel, and completing other multistep tasks. TechRepublic previously reported that Muse launched in the U.S. this month across iOS, Android, the web, and WhatsApp. Meta Business Agent is aimed more directly at companies. Meta says more than one million businesses already use a Business Agent on WhatsApp and Messenger to respond to customers. The agent can answer questions, recommend products, book appointments, qualify leads, and close sales. Meta also says more than one billion active threads with businesses take place across WhatsApp, Messenger, and Instagram each day. For larger organizations, Meta Business Agent Platform can connect with hundreds of systems, including Shopify, Zendesk, and Shopee. Meta says the platform includes enterprise controls, guardrails, and measurement tools that let companies define how their agents operate. Meta has also outlined broader ambitions for Business Agent. The company says it eventually wants the technology to help with tasks such as market research, product insights, calendar management, and competitive intelligence. Enterprise Platform now gives those efforts a dedicated business operation. Meta hires an enterprise software veteran to lead the push Meta recruited former MongoDB CEO Chirantan "CJ" Desai as chief enterprise platform officer. He will report directly to Zuckerberg. Before leading MongoDB, Desai oversaw product and engineering at Cloudflare and spent nearly eight years at ServiceNow, including as president and chief operating officer. "Meta Enterprise Platform will focus on turning its AI stack into products and services that companies can deploy for their own businesses," Desai said in Meta's announcement. The hire gives Meta an executive with experience building and selling software to large organizations as it moves beyond its traditional advertising and consumer-platform businesses. Microsoft and Salesforce have another AI competitor Meta is entering a market where major enterprise software companies have already spent years embedding AI into existing business platforms. Microsoft offers AI agents across Microsoft 365 Copilot and Dynamics 365, including agents that can access organizational data and take actions such as sending emails or updating records. Salesforce has similarly expanded Agentforce across its CRM products, positioning autonomous agents around sales, service, marketing, and other customer workflows. Meta starts from a different position. It does not have Microsoft's productivity suite or Salesforce's CRM footprint, but it does have massive consumer reach, business messaging platforms, AI models, agents, and large-scale infrastructure. That could give Meta an opening through customer-facing operations before it pushes further into internal enterprise workflows. The move could also provide another way for Meta to generate returns from its rapidly expanding AI infrastructure. According to Meta, the company currently expects 2026 capital expenditures of between $130 billion and $145 billion, with data center capacity among the drivers of that spending. Zuckerberg has previously said Meta sees a broader enterprise opportunity that could include APIs, business agents, services for large customers, and potentially selling compute directly. Meta still has big enterprise questions to answer For all the ambition behind the announcement, Meta Enterprise Platform is still more of an enterprise strategy than a clearly packaged software suite. Meta has not disclosed overall platform pricing, named launch customers, or explained exactly how Muse, Business Agent, Muse API, and Muse Code will be packaged and administered together. Those details will matter for IT leaders deciding whether Meta belongs alongside established enterprise vendors. Organizations will need to understand how administrators control agent access, which actions require human approval, how activity is logged and audited, and how Meta's products integrate with existing identity, security, and business systems. Meta says security and privacy will be built into its enterprise products from the outset. Its existing Business Agent Platform already includes enterprise controls and guardrails, but Meta has not yet provided the same level of detail for the broader Enterprise Platform. That leaves the biggest question unanswered: whether Meta can turn its enormous consumer reach and AI investment into an enterprise platform companies trust with core business operations. For now, Meta is no longer content with helping businesses reach customers through its apps. It wants its AI agents doing some of the work behind them, too. Editor’s note: This article originally appeared on our sister publication, TechRepublic. More
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
By DZone Staff
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
By Jubin Soni, FBCS DZone Core CORE
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
By Uthej Mopathi DZone Core CORE
How to Test POST API Requests With Playwright TypeScript
How to Test POST API Requests With Playwright TypeScript

Testing POST API requests is an important skill for modern QA and automation engineers working with backend services and microservices. In this article, we’ll explore how to test POST API requests with Playwright and TypeScript, focusing on sending a request body using different approaches. By the end, you will learn how to send POST API requests using the following approaches for adding a request body: JSON Object/Array(data)Stringified JSONJSON FileFaker library Application Under Test We will use the POST /addOrder API of the RESTful e-commerce demo application for this demo. The API schema is provided below: JSON { "user_id": "string", "product_id": "string", "product_name": "string", "product_amount": 0, "qty": 0, "tax_amt": 0, "total_amt": 0 } ] Testing POST API Requests With Playwright TypeScript Playwright provides a powerful request API that allows us to create and manage HTTP request contexts. Let’s walk through how to send POST API requests step-by-step using different approaches for passing the request body: Request Body as JSON Object/Array Let’s send a POST API request with a static JSON Array and verify that the response status code is 201. TypeScript test("POST order details API with static JSON Array", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data:[{ user_id: "1", product_id: "82", product_name: "Cadbury Bar", product_amount: 12, qty: 2, tax_amt: 1, total_amt: 25 }, { user_id: "2", product_id: "80", product_name: "MilkyBar", product_amount: 10, qty: 1, tax_amt: 1, total_amt: 11 } ], }); expect(response.status()).toBe(201); }); Code Walkthrough A Playwright test case for testing the POST API request is created using the request API, which is a built-in Playwright APIRequestContext fixture used to send HTTP requests. Sending the POST Request: The request.post() sends a POST request to the “http://localhost:3004/addOrder/” endpoint. The response received after sending a POST request is stored in the response variable.Passing the Request Body (Static JSON Array): The data property is used to send the request body, which accepts data in JSON Array and JSON object formats. Since the POST /addOrder API accepts an array of order details, allowing multiple orders to be submitted within a single JSON array, the test provides the data in JSON array format.Validating the Response: The response.status() retrieves the HTTP status code. The assertion ensures that the status code 201 is returned, indicating that the orders were successfully created. Request Body Using JSON.Stringify() JSON.stringify() can be used while sending raw payloads, testing malformed JSON, or sending custom-formatted JSON. However, the Content-Type header should be set to application/json so the server correctly interprets the request body as JSON data. Let’s write the same POST /addOrder test using JSON.stringify() and verify that a status code 201 is returned in the response. TypeScript test("POST order details API using JSON.Stringify", async ({ request }) => { const orderData = [{ user_id: "5", product_id: "64", product_name: "Cadbury Mini", product_amount: 5, qty: 3, tax_amt: 1, total_amt: 16 }]; const response = await request.post("http://localhost:3004/addOrder/", { data: JSON.stringify(orderData), headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); The orderData array is defined to contain one order object. Each field represents the order details such as user_id, product_id, product_name, and so on. This is a normal JavaScript object at this stage, not JSON yet. The JSON.stringify(orderData )converts the JavaScript object into a JSON string. As we manually stringify the request payload, we should explicitly set the Content-Type header to application/json. It tells the server to treat the request body as JSON data. Playwright sends the POST API request using the post() method and then validates that the response returns a 201 status code. Request Body as a JSON File Using a JSON file as a request body is a handy approach when testing POST API requests with Playwright TypeScript. It supports large payloads and allows the same file to be easily reused across multiple tests. The following orders.json file will be used as a request payload in the POST /addOrder API: JSON [ { "user_id": "1", "product_id": "79", "product_name": "5 star 10gm Chocobar", "product_amount": 5, "qty": 1, "tax_amt": 0.5, "total_amt": 5.5 }, { "user_id": "2", "product_id": "71", "product_name": "Lindt Milk Chocolate", "product_amount": 15, "qty": 3, "tax_amt": 2.5, "total_amt": 47.5 } ] The following configurations should be in place before we proceed to write the test using a JSON file as the request payload. tsconfig.json JSON { "compilerOptions": { "module": "NodeNext", "moduleResolution": "NodeNext", "resolveJsonModule": true } } We need to ensure that the JSON file is imported first and then attached to the data parameter while sending the POST request. TypeScript import orders from '../test_data/orders.json' with {type: 'json'}; test("POST order details API using JSON file", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data: orders, }); expect(response.status()).toBe(201); }); This test imports test data from an external orders.json file and uses it as the request body for a POST API call in Playwright. The request.post() method sends the imported JSON directly as the payload, and the test verifies that the API returns a 201 status code. Using a JSON file makes it easier to manage, reuse, and update request payloads without modifying the test logic, improving the maintainability and readability of API tests. Request Body With Faker Library Using the Faker library to create a request body for a POST API helps generate realistic and dynamic test data, reduces hardcoded values, and improves test coverage. It is especially useful for simulating real-world scenarios and avoiding duplicate data issues. However, it also has limitations. Since Faker generates random data, it may lead to inconsistent test results if not properly handled, and debugging failures could become more difficult without fixed or reproducible inputs. To use the Faker library, we need to install it first using the following command: Plain Text npm install --save-dev @faker-js/faker Next, let’s create two helper functions: one to create order objects with the required order details, and another to generate an array of orders based on the count provided by the user. TypeScript import { faker } from '@faker-js/faker'; export function createOrderDetails() { const productAmount:number = faker.number.int({ min: 1, max: 100 }); const qty:number = faker.number.int({ min: 1, max: 5 }); const taxAmt:number = faker.number.int({ min: 2, max: 10 }); const totalAmt:number = (productAmount*qty)+taxAmt; return { user_id: faker.number.int({ min: 1, max: 50 }), product_id: faker.number.int({ min: 1, max: 100 }), product_name: faker.commerce.productName(), product_amount: productAmount, qty:qty, tax_amt: taxAmt, total_amt: totalAmt }; } The createOrderDetails() function returns the order details with the required fields. It also calculates the total amount by summing the total product value and tax amount. The values in all fields are updated randomly using the appropriate methods provided by the Faker library. TypeScript import { faker } from '@faker-js/faker'; export function createRandomOrders(count:number) { return faker.helpers.multiple(createOrderDetails, {count}); } The createRandomOrders(count) method accepts a count parameter and returns randomly generated orders that can be directly used in the test. TypeScript return faker.helpers.multiple(createOrderDetails, {count}); This is where the work happens. multiple() is a helper function provided by Faker. Its job is to execute another function multiple times and collect all the results into an array. It takes two arguments: A function to execute repeatedly.An options object specifying how many times to execute it. It asks Faker to call createOrderDetails() exactly count times, collect all the generated orders into an array, and return that array to the caller. TypeScript test('POST order details API using Faker library', async({request}) => { const orderData = createRandomOrders(5); const response = await request.post("http://localhost:3004/addOrder/", { data: orderData, headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); This test sends a POST request to the /addOrder endpoint API using dynamically generated test data from the Faker library. The createRandomOrders(5) function generates an array of 5 random order objects, which are passed as the request body using the data property. The test then verifies that the API responds with a 201 status code, confirming that the orders were successfully created. Summary In this tutorial, we explored multiple approaches to testing POST API requests using Playwright with TypeScript, including sending static JSON payloads, using JSON.stringify, importing data from external JSON files, and generating dynamic test data with the Faker library. We also covered how to handle headers correctly and validate API responses using status code assertions. In my experience, using external JSON files is a practical approach for testing POST API requests because it lets us send bulk data from a single file. Similarly, the Faker library can also be used to generate dynamic test data; however, in some cases, using a third-party library may not be permitted. Ultimately, the software team chooses the approach that best fits their testing strategy. Happy testing!

By Faisal Khatri DZone Core CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1

A few months ago, one of our teams celebrated a milestone in their quarterly review. AI adoption was up. The productivity dashboard showed that developers were generating, on average, 40% more code per sprint. The tech lead showed this as a major win. Three weeks later, I was on a call for a production issue. The challenge was that errors were coming in three different formats depending on which endpoint you hit. The alerting was blind to a category of failure it had always caught before. We had a centralized exception handler. It logged context, mapped it to the right HTTP status, and pushed these alerts to our observability stack. When we investigated, we found that recent AI-assisted PRs had started introducing their own try-catch blocks inline. Each one caught exceptions locally, logged in a slightly different format, and returned a slightly different error shape. Some swallowed the exception instead of letting it propagate up to the handler that would have alerted us. Each one of those PRs was correct. Every one passed review, including the reviews I did myself. We were all checking for correctness, and consistency isn't the kind of thing that shows up in a git diff. Cleaning it up took most of a sprint. And that productivity dashboard? It counted the original generation and the cleanup as output. Twice the code, twice the "productivity," for a net loss of engineering time. The dashboard was going up while the system got worse underneath it. That's the mistake I keep seeing where teams confuse code generation with engineering progress. The LOC Trap, Reloaded Fred Brooks called this out in The Mythical Man-Month decades ago. He explicitly mentioned that measuring programming productivity by lines of code is nonsensical. Everyone agreed and then somehow forgot. So why did we rebuild the exact same dashboard the moment AI arrived? We are using the same flawed metric, with AI branding, and presenting it to boards. When a team lead reports that AI tools helped produce 40% more code, the follow-up I want to hear is “Did we actually need 40% more code?” Usually the answer is no. What we needed was the same outcomes with less effort, and effort in software lives overwhelmingly outside the act of typing. I’ll come back to that. There is a difference this time, and it is worth naming. Teams aren’t defending line counts out loud anymore. The dashboards have moved on to merged pull requests, agent tasks completed, and suggestions accepted. It is the same instinct in a unit that sounds more respectable in front of a board, counting the artifacts of work and reporting the count as progress. Why Senior Engineers Delete Code Here’s a pattern you’ll recognize if you have led engineering teams for any length of time. Your best engineers often produce fewer lines of code than anyone else on the team. Some of their most impactful weeks come out to negative line counts. That’s expertise rather than laziness. I haven’t fully worked out why the instinct for deletion over addition takes years to develop, but it does. GitClear has been measuring this rather than speculating about it. Their 2026 analysis covers 623 million changes from 2023 to 2026, and the figure that stopped me had nothing to do with volume. Moved code, their proxy for refactoring, dropped from 21% of all changes in 2022 to 3.8% by the middle of this year. Duplicated blocks are up 81% across the same window. Cross-file function calls, which is roughly what reuse looks like in a diff, are down 35%. Those three together describe a codebase that has stopped being rearranged when work goes in, and very little gets moved or deleted. When someone takes 2,000 lines of tangled logic and replaces it with 200 clean ones, that looks like a loss on any volume metric. It is an enormous win for the system, and it is exactly the activity that has gone quiet. One correction I owe, since I quoted the earlier version of this research at people for the better part of a year. GitClear’s 2024 report predicted two-week code churn would double in the AI era. It didn’t double. It went up 15%. The headline projection was too aggressive, and the part almost nobody quoted — refactoring falling off a cliff — turned out worse than predicted. Senior engineers get this intuitively. Every line of code is a liability, because every line has to be read, understood, tested, and maintained. So the best solution often makes code disappear, and that takes different forms like a well-chosen abstraction that kills duplication or a config change that removes a custom implementation. Sometimes it’s just a conversation with the PO that drops the requirement entirely. Now think about what AI coding metrics would say about this. An engineer spends a day understanding a system, realizes three services can collapse into one, and deletes 4,000 lines. By every AI productivity metric in use today, that engineer had a terrible day. In reality, they may have saved the organization months of future pain. Gergely Orosz tells a revealing story about what happens when you optimize for the wrong signal. When Uber introduced diff-count metrics, engineers started creating more, smaller changes to look productive. They flooded CI systems, driving up costs. The metric improved and engineering got worse. We are setting ourselves up for the same trap with AI-generated LOC. AI’s Tendency Toward Verbose Implementations This gets worse when you look at what AI coding tools actually excel at, which is producing plausible code quickly. That skill carries a built-in bias toward verbosity. Sometimes more code is genuinely the right call. Explicit beats implicit. A verbose but readable implementation can be better than a clever one-liner that nobody understands at 3 am when production is on fire. I’m not arguing for code golf. But AI-generated verbosity is a specific kind of bad, because it is default verbosity from ignorance of context rather than chosen verbosity for clarity. This distinction matters a lot more than I initially thought. Ask an AI assistant to implement a feature, and you’ll get a complete, working solution, longer than what an experienced developer would write, because it optimizes for correctness and completeness in isolation. It may not pick up the utility you wrote last month to do exactly this. It doesn’t realize the framework provides a one-liner if you structure the problem slightly differently. It can’t tell the difference between “I should be explicit here for readability” and “I am reinventing something that already exists three directories over.” The handler drift I opened with is the cleanest example I have of it. The AI did exactly what it was asked, every single time. Each PR added code, each one passed review on its own terms, and nothing was wrong inside any of them. What we lost lived across them. One handler gave us one error shape, and one error shape gave our alerting something to fire on. No single diff broke that, and all of them together did. GitClear tracks error-masking constructs, which is the failure mode buried in there, and they are up 47% since 2023. Our inline handlers were exactly that. I’d like to think we were an unlucky outlier, and the data says we were ordinary. I’m still not sure how you review for a property that isn’t visible in the file in front of you. Complexity as the Hidden Cost Code volume isn’t a perfect proxy for system complexity, and I acknowledged that above. But it is a directional one, and in aggregate it holds. When your codebase grows by 30-40% in a quarter without a corresponding growth in functionality, complexity is almost certainly growing with it. And system complexity is the single biggest thing determining how fast your team can move over time. I haven’t found a way around that in eighteen years. Every line of code carries ongoing costs that nobody puts on a dashboard: Cognitive load for anyone working nearbyTest coverage, without which it becomes a ticking riskReview time on every future change to itDependencies it drags in that need constant updatingMigration effort during every platform change When AI tools grow your code volume by 30-40%, each of these costs grows. The productivity gain at the moment of writing is real, and I’m not denying that. But it can be entirely eaten up by the downstream cost of maintaining a bigger, more complex system. Then there is the study I keep coming back to, and the update to it that I nearly missed. In mid-2025, METR found that experienced developers using AI coding tools took 19% longer on real-world tasks. The setup was specific. It had 16 open-source contributors working in their own repositories, code they’d lived in for years, 246 real issues averaging about two hours each, with Cursor Pro and Claude. In February 2026, METR published an update that takes a fair amount of it back. They are redesigning the experiment, and the reasons aren’t flattering to the original. Developers who would no longer work without AI declined to take part at all. Somewhere between 30% and 50% of participants avoided submitting exactly the tasks where they expected to want AI. The pay rate for the follow-up work dropped from $150 an hour to $50, which made recruitment worse again. Their own summary is that the newer data amounts to very weak evidence in either direction, and that developers are probably more sped up now, in early 2026, than the early-2025 estimate suggested. So the 19% was never a fact about AI-assisted development. It was a measurement of sixteen people in one setting, and the people who ran it now think it read low. What survives is the part that makes me think harder. Those developers believed they were 20% faster. Whatever the true effect was, it wasn’t the effect they perceived, and not one of them could feel the gap while it was happening. Reviewing and integrating suggestions consumed time none of them accounted for. Selection bias moved the headline number around, but it does not explain away a room full of experienced engineers being wrong about their own week. I could never square that finding with my own experience, because I do feel faster on certain tasks. The update moves me off the fence, slightly. Maybe I wasn’t fooling myself. I would still put no weight at all on my own estimate of how much faster I am, and that’s close to the only thing here I am confident about. None of that is an argument against the tools, only against measuring them by how much code they produce. The Strongest Number Against Me If you want to argue the other side, the best evidence available today is Microsoft’s. Early in 2026, they rolled Claude Code and GitHub Copilot CLI out across the organization and studied what happened. Tens of thousands of engineers, four months, and the ones who adopted merged roughly 24% more pull requests than the counterfactual said they would have. That is not a lab, and it is not a small group of volunteers. It’s the largest measurement of agentic coding tools anyone has published, the effect is large, and it points the right way. I take it seriously. I also notice what the unit is. The authors get there ahead of any critic. Their paper says a merged PR isn’t the same as the value it delivers, which is the argument of this entire post, conceded inside the study that is supposed to answer it. Twenty-four percent more merged pull requests is consistent with 24% more delivered value. It is equally consistent with the same work arriving in smaller slices, which is what happened at Uber the moment diff count landed on a dashboard. There are narrower caveats, and I won’t pretend I have chased all of them. The comparison is against engineers who already had AI in their IDE, so what it measures is the increment from adding an agent rather than the effect of AI from zero. Four months is not long enough for maintenance cost to turn up. And engineers chose for themselves whether to adopt. None of that makes the study wrong. It is a good measurement of pull request volume, and pull request volume behaves the way lines of code always did. It is easy to report and showcase, if somebody decides that moving it matters. Blind to what the system underneath is doing. What This Actually Means If You Are Leading a Team If you’re six months into your AI investment and your main evidence of ROI is that the output counter went up, whether that counter says lines or pull requests or tasks completed, that should worry you more than it reassures you. You might be measuring the accumulation of future cost and calling it present-day value. Here is what I would look at instead. Cycle time as value reaching production sooner, not just code getting written faster? Those two come apart more often than anyone expects. Rework rate counts if you are fixing more bugs in AI-assisted code? If generated code carries a higher defect rate, your productivity gain is a mirage. Cognitive complexity trend is when your application is getting harder to understand. Tools like SonarQube measure this. If complexity is climbing faster than it did before AI adoption, you’ve got a compounding problem. Developer effort distribution is where the time is actually going. If writing dropped from 20% to 10% of developer time while code review grew from 15% to 30%, you have moved the burden rather than reduced it. None of that is exotic, and it is not just my read. DORA’s 2026 work on the ROI of AI-assisted development arrives at something similar from a different direction. The return on these tools tracks the strength of the engineering system around them rather than the tools themselves. The things that decide it are unglamorous. Whether code review has any slack left in it. Whether people trust the test suite enough to act on a red build, which is the one I have never seen anybody audit. How much sits between a merge and production. Where those are weak, faster generation fills a queue and waits there, and the measurement science is still catching up to the tooling. The Uncomfortable Question Here is what I would ask any engineering leader who reports AI productivity gains based on code output: “If your best engineer spent last week deleting 3,000 lines of AI-generated code and replacing them with 300 lines that do the same thing better, would your dashboard show that as a win or a loss?” On a line count, that week is a catastrophe. On a PR count, it is one merged pull request, which is what a typo fix is worth too. Neither number has any way of seeing what actually happened. If your dashboard shows a loss, you are measuring the wrong thing, and you are rewarding your team for building a larger, slower, more fragile system in exchange for a chart that goes up and to the right. The goal was never more code. It was better systems that deliver business value and are maintained cheaply. Next in this series: Optimizing Benchmark Tasks Instead of Real Delivery Work — why the fact that coding is not the bottleneck makes most AI productivity claims irrelevant to actual delivery speed.

By Gaurav Gaur DZone Core CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads

Batch processing remains vital because many business operations aren't suited to interactive requests. Tasks such as recalculating prices, reconciling transactions, migrating records, generating reports, processing invoices, reclassifying customers, or applying rules across millions of records may require considerable time. Handling these as standard requests leads to fragile systems, increased user wait times, frequent timeouts, challenging retries, and possible data inconsistencies. A batch model handles large workloads predictably, incrementally, and with control over progress and recovery. Rather than processing a massive operation as a single loop, batch processing uses jobs, steps, chunks, checkpoints, filtering, and restartability. This approach separates long-running data tasks from the user experience while delivering a structured execution model. In this article, we will focus on Jakarta Batch and examine its sustained relevance for modern enterprise applications. Why Batch Processing Still Matters in Enterprise Systems Modern applications offer various methods for background processing, such as message queues, event-driven architectures, schedulers, reactive pipelines, and distributed stream-processing platforms. While each addresses specific needs, batch processing is most effective when operations have a defined start and end, involve a known or discoverable dataset, and require controlled execution, progress tracking, restartability, or periodic processing. Batch processing remains essential in enterprise systems. Workloads such as financial reconciliation, billing, payroll, reporting, data migration, regulatory processing, catalog updates, and large-scale reclassification are still prevalent. In these scenarios, the priority is to process large volumes of work safely and predictably, rather than responding to individual events quickly. Batch provides a model specifically designed for these requirements. How Jakarta Batch Works Jakarta Batch organizes background processing into jobs and steps. A job defines the overall batch operation, while each step represents a specific stage. In chunk-oriented processing, a step follows a simple pipeline: read, process, write, and repeat until it processes all input. The Jakarta Batch runtime manages this lifecycle so application code can focus on reading, transforming, and persisting data. A job is the top-level unit of execution and represents a complete business operation, such as importing records, recalculating customer classifications, processing invoices, or reconciling transactions. Jobs can accept parameters at startup, allowing the same batch definition to run with different inputs or business rules. A job consists of one or more steps, each representing a distinct phase of the workload. Simple jobs may have a single step, while complex processes can use multiple steps in sequence, such as importing data, validating it, and generating a final report. Within a chunk-oriented step, the ItemReader supplies data to the runtime one item at a time, from sources such as a database or file. The reader only retrieves the next item and does not need to know how it will be processed or persisted.The ItemProcessor receives each item and applies business rules, such as validation, transformation, classification, enrichment, or filtering. It may return a modified item or null if the item should be excluded from writing.The ItemWriter receives processed items and persists or exports them. Unlike the reader and processor, which handle items individually, the writer typically receives a group of items from the current chunk. This enables more efficient database or bulk operations. Jakarta Batch adds features around this pipeline to support enterprise workloads. The runtime manages chunk boundaries, transactions, checkpoints, execution status, failures, and restart behavior. Chunk size determines how much work is grouped before a write and checkpoint, making it a key parameter for balancing throughput, memory usage, database cost, and recovery. The core model is straightforward: Job → Step → Read → Process → Write → Repeat Jakarta Batch keeps the business pipeline simple while the runtime manages the execution mechanics needed for reliable, long-running data processing. The Sample: Customer Segmentation with Jakarta Batch This example demonstrates the Jakarta Batch model using an e-commerce customer segmentation scenario. Customers are assigned to tiers such as Bronze, Silver, Gold, and Platinum based on configurable spending thresholds. When thresholds change, the application reevaluates the customer base and updates only customers whose classification has changed. The full application includes MongoDB integration, a Jakarta Faces UI, a preview workflow, validation, and supporting services. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. This section focuses on the classes directly involved in Jakarta Batch execution. Starting the Batch Job The application initiates the batch process through CustomerSegmentationService. Unlike the reader, processor, and writer, this class is not a batch artifact. Instead, it is an application service that retrieves Jakarta Batch’s JobOperator from BatchRuntime to start and monitor job executions. Java @ApplicationScoped public class CustomerSegmentationService { public static final String JOB_NAME = "customer-segmentation"; private volatile CustomerSegmentationPolicy currentPolicy; // initialization and status methods omitted public long start(CustomerSegmentationPolicy policy) { if (isRunning()) { throw new IllegalStateException( "A customer segmentation batch is already running"); } Properties parameters = new Properties(); parameters.setProperty( CustomerSegmentationPolicy.JOB_PARAMETER, policy.toJson()); long executionId = BatchRuntime.getJobOperator() .start(JOB_NAME, parameters); currentPolicy = policy; return executionId; } public boolean isRunning() { // implementation omitted } } The key API here is JobOperator, which Jakarta Batch provides as the interface for starting, stopping, restarting, and inspecting jobs. In this example, the segmentation policy is serialized into the job parameters to ensure each execution gets the correct business rules. Reading the Input The first batch artifact, CustomerItemReader, extends Jakarta Batch’s AbstractItemReader to implement a chunk-oriented reader. Java @Named("customerItemReader") @Dependent public class CustomerItemReader extends AbstractItemReader { @Inject private CustomerRepository customerRepository; private List<Customer> customers = List.of(); private int nextIndex; @Override public void open(Serializable checkpoint) { try (Stream<Customer> customerStream = customerRepository.findAll()) { customers = customerStream .sorted(Comparator.comparing(Customer::getId)) .toList(); } nextIndex = checkpoint instanceof Integer index ? index : 0; } @Override public Customer readItem() { if (nextIndex >= customers.size()) { return null; } return customers.get(nextIndex++); } @Override public Serializable checkpointInfo() { return nextIndex; } } These methods are part of the Jakarta Batch reader lifecycle defined by AbstractItemReader. open() prepares the reader and accepts a previous checkpoint if available. readItem() provides the next item to the runtime; returning null indicates there is no more input. checkpointInfo() reports the reader’s current position for checkpointing. For simplicity, this sample loads customers into memory. For larger workloads, the implementation might use pagination or a MongoDB cursor without changing the Jakarta Batch model. Processing Each Customer The next artifact implements Jakarta Batch’s ItemProcessor interface. Java @Named("customerTierProcessor") @Dependent public class CustomerTierProcessor implements ItemProcessor { @Inject @BatchProperty( name = CustomerSegmentationPolicy.JOB_PARAMETER) private String thresholdsJson; private CustomerSegmentationPolicy policy; @PostConstruct void initialize() { policy = CustomerSegmentationPolicy.fromJson( thresholdsJson); } @Override public Customer processItem(Object item) { if (!(item instanceof Customer customer)) { throw new IllegalArgumentException( "Expected a Customer item"); } CustomerTier calculatedTier = policy.tierFor(customer.getTotalSpent()); if (calculatedTier == customer.getTier()) { return null; } return Customer.builder() .id(customer.getId()) .name(customer.getName()) .totalSpent(customer.getTotalSpent()) .tier(calculatedTier) .build(); } } Here the Jakarta Batch contract is explicit: ItemProcessor defines processItem(). The runtime calls that method for every item produced by the reader. The processor applies the segmentation rule and either returns the transformed customer or null. Returning null has a specific meaning in Jakarta Batch: the item is filtered and does not continue to the writer. The @BatchProperty is also part of the Batch integration. It receives the thresholds property defined for this job execution, allowing the processor to reconstruct the CustomerSegmentationPolicy before processing begins. Writing the Results The final artifact extends AbstractItemWriter, Jakarta Batch’s base implementation for writing a chunk. Java @Named("customerItemWriter") @Dependent public class CustomerItemWriter extends AbstractItemWriter { @Inject private CustomerRepository customerRepository; @Override public void writeItems(List<Object> items) { List<Customer> customers = items.stream() .map(this::toCustomer) .toList(); customerRepository.saveAll(customers); } private Customer toCustomer(Object item) { if (item instanceof Customer customer) { return customer; } throw new IllegalArgumentException( "Expected a Customer item"); } } writeItems() is defined by the Jakarta Batch writer contract inherited from AbstractItemWriter. Unlike the processor, which receives one item at a time, the writer receives a collection of processed items. In this case, the collection contains only customers whose classification changed, as the processor has already filtered the others. At this point, the Java components of the pipeline are as follows: Plain Text CustomerItemReader extends AbstractItemReader ↓ CustomerTierProcessor implements ItemProcessor ↓ CustomerItemWriter extends AbstractItemWriter These types are what connect the application code to the Jakarta Batch runtime. Connecting the Artifacts With JSL The Java classes define the behavior, but Jakarta Batch requires explicit mapping of the reader, processor, and writer to each job. This orchestration is described in JSL: XML <?xml version="1.0" encoding="UTF-8"?> <job id="customer-segmentation" xmlns="https://jakarta.ee/xml/ns/jakartaee" version="2.0"> <step id="recalculate-customer-tiers"> <chunk item-count="20"> <reader ref="customerItemReader"/> <processor ref="customerTierProcessor"> <properties> <property name="thresholds" value="#{jobParameters['thresholds']}"/> </properties> </processor> <writer ref="customerItemWriter"/> </chunk> </step> </job> The ref values correspond directly to the names declared with @Named in the Java classes: Java @Named("customerItemReader") @Named("customerTierProcessor") @Named("customerItemWriter") The XML therefore tells the Jakarta Batch runtime: for this step, use this reader, then this processor, and finally this writer. It also maps the thresholds job parameter into the processor property. The item-count="20" sets the chunk size for this sample. Jakarta Batch coordinates reading and processing, periodically invoking the writer according to the chunk lifecycle and establishing transaction and checkpoint boundaries. The value 20 is for demonstration; real applications should tune chunk size based on processing cost, database behavior, transaction size, throughput, and recovery requirements. This structure is recommended for the article: present the class declaration first, then describe the lifecycle methods inherited from or required by Jakarta Batch. This approach helps the sample teach the API rather than simply presenting isolated methods. Conclusion Jakarta Batch is valuable because it transforms large-scale data processing into a structured execution model, eliminating the need for custom loops and ad hoc background logic. By separating reading, processing, and writing, and introducing runtime concepts such as jobs, steps, checkpoints, restartability, and chunk-oriented execution, it provides enterprise applications with a predictable approach to handling workloads involving thousands or millions of records. This allows implementations to focus on business logic, while the Batch runtime manages repetitive execution concerns, making the model easier to understand, optimize, and scale as workloads increase.

By Otavio Santana DZone Core CORE
The Warning That Never Stops the Agent
The Warning That Never Stops the Agent

I went looking for a flag that could cap the amount spent on a session on DeepAgents. The kind of hard spending cap that stops an agent loop before it burns through a budget. What exists instead is a warning, and the gap between "warning" and "stopping" turned out to be the more interesting story. What Already Exists: A Number With No Teeth deepagents-code has a CostTrackingMiddleware that owns a thread's cumulative spend. This is a real, checkpointed dollar figure priced from actual per-request token usage, not an estimate. A separate feature built on top of that number, merged earlier, adds a one-time warning once the total crosses a configured threshold (default $50): Python # app.py, roughly what ships today threshold = self._session_cost_warning_threshold_usd if ( not self._session_cost_warning_shown and 0 < threshold < self._session_cost_usd ): self._session_cost_warning_shown = True self.notify( f"Estimated session cost is {format_cost(self._session_cost_usd)}, " f"above the configured {format_cost(threshold)} threshold. Consider " "/offload to reduce context usage or /clear to start fresh.", title="Session cost warning", severity="warning", timeout=12, markup=False, ) Read that once more: It's a self.notify(...) call. A toast! The agent's own loop has no idea this happened. Nothing in the code above touches the graph, the model call, or the next tool invocation. If you're watching the terminal, you see the warning and can intervene by hand (/offload, /clear, or just killing the process). If you're not watching and it's a headless CI run — for example, an unattended overnight session or a tool-call loop that's quietly retrying against a flaky API — this number keeps climbing, and nothing stops it. That's the gap: a cost number that can only ever inform a human, never the loop that's actually spending the money. Why the Fix Isn't "Just Check the Number Somewhere" The obvious instinct is to add a if cost > limit: stop check. The real work is in where that check has to live and what "stop" has to mean to a LangGraph agent loop. Where: The check needs to run before the next model call, not after. Checking after a call has already happened is too late to prevent its cost. LangGraph gives middleware a before_model hook for exactly this. LangChain's own ModelCallLimitMiddleware (a call-count limiter, not a cost one) already establishes the pattern: check a condition in before_model, and if it's tripped, return an update that redirects the graph to end instead of letting the model call happen. What "stop" means: A plain return None from before_model just lets the loop continue. To actually halt, the hook needs @hook_config(can_jump_to=["end"]) and has to return {"jump_to": "end", ...} which is a real graph-control signal, not a value the caller has to notice and act on: Python # cost_tracking.py — the actual hook, as committed @hook_config(can_jump_to=["end"]) def before_model( self, state: CostState, runtime: Runtime[ContextT], ) -> dict[str, Any] | None: """Halt the run before the next model call if the hard cost cap is met. Checked against the checkpointed cumulative total from the *previous* step -- the same figure the TUI's soft warning reads via `_set_session_cost` -- so this fires at the same point in the loop a user would already have seen the warning toast, just before the next request that would push spend further over the configured cap. """ if self._nested or self._hard_limit_usd is None: return None total_usd = state.get("_session_cost_usd") if ( isinstance(total_usd, bool) or not isinstance(total_usd, int | float) or not math.isfinite(total_usd) ): return None if total_usd < self._hard_limit_usd: return None if self._exit_behavior == "error": raise CostLimitExceededError(total_usd, self._hard_limit_usd) limit_message = _build_cost_limit_message(total_usd, self._hard_limit_usd) return {"jump_to": "end", "messages": [AIMessage(content=limit_message)]} Two details worth calling out, because they're the kind of thing that looks like overcaution until you hit the failure it's guarding against: isinstance(total_usd, bool) before the numeric check. In Python, bool is a subclass of int, so isinstance(True, int | float) is True and True < 5.0 evaluates fine (True == 1). Without the explicit bool guard, a stray True sitting in a state field meant for a float would silently be treated as $1.00 and could either falsely trip the halt or falsely pass it, depending on the cap. Cheap to guard against, expensive to debug if you don't. Checked only on the non-nested instance. CostTrackingMiddleware also runs on subagents, where its _session_cost_usd channel tracks that subagent's own local spend before it's transferred back to the parent's running total. Checking the hard cap there would be checking the wrong number - a subagent's small local total against a cap meant to bound the whole session's spend. self._nested gates this out entirely. The Part That Isn't the Algorithm: Getting the Number Across a Process Boundary Here's what made this bigger than a one-file change: dcode doesn't run the agent loop in the same process as the CLI. It starts a LangGraph server in a subprocess and talks to it over langgraph-sdk. Every configuration value the agent needs, like model name, sandbox type, recursion limit, and now this cap, has to survive that boundary, which in this codebase means round-tripping through environment variables on a ServerConfig dataclass: Python # _server_config.py max_cost_usd: float | None = None """Explicit hard cap, in USD, on the main thread's cumulative estimated cost.""" def to_env(self) -> dict[str, str | None]: return { ... "MAX_COST_USD": ( str(self.max_cost_usd) if self.max_cost_usd is not None else None ), ... } @classmethod def from_env(cls) -> ServerConfig: return cls( ... max_cost_usd=_read_env_float("MAX_COST_USD", default=None), ... ) _read_env_float didn't exist before this - every other numeric config value in this file is an int (recursion_limit, turn counts), so there was no float-reading helper to reuse. One new function, matching the existing _read_env_int's shape exactly, and the round-trip works the same way every other config value already does. The full path, in order: a --max-cost CLI flag → a resolver that checks the flag, then a config.toml entry, then "disabled" → create_cli_agent(max_cost_usd=...) → ServerConfig → serialized to an environment variable → the subprocess reads it back → create_cli_agent again, this time inside the subprocess → CostTrackingMiddleware(hard_limit_usd=...). Six hops for one float, and every one of them was necessary - skip the ServerConfig round-trip and the flag works in a unit test but silently does nothing the moment you actually run dcode, because the subprocess that runs the real agent loop never sees it. Verifying It Live Unit tests covering the before_model logic in isolation are necessary but not sufficient here. They'd pass even if one of those six hops silently dropped the value, because a unit test calls the middleware directly and never exercises the subprocess boundary at all. The only way to know the flag actually works is to run the real CLI against a real model: Shell $ dcode -n "Read sample.txt, then read it again, then read it a third time. \ Do this as three separate tool calls, one per turn." \ --model anthropic:claude-haiku-4-5 \ --max-cost 0.0001 \ --max-turns 6 I'll read sample.txt, then read it again on the next turn, then a third time on the turn after that. First read: Calling tool: read_file Session halted: estimated cost $0.02 has reached the configured limit of <$0.01. Raise the limit (e.g. `--max-cost`) or start a new session to continue. Task completed Usage Stats Provider Model Reqs InputTok OutputTok Cost anthropic claude-haiku-4-5 1 13.6K 147 $0.02 The model was told to make three tool calls, one per turn. It made exactly one. The middleware checkpointed that turn's real cost (0.02), the next beforemodel check found the cumulative total over the(deliberately absurd) 0.0001 cap, and the graph jumped to end with the injected message instead of continuing to spend on turns two and three. Total cost of proving this worked: two cents. Why This Is Worth a Hard Stop and Not Just a Bigger Warning You could imagine closing this gap by making the warning louder: repeat it every turn instead of once, or block user input until it's acknowledged. That doesn't fix the actual failure mode, which is specifically the unattended case. A louder toast is still a toast. Nothing short of a return value the graph itself has to obey closes that gap, which is why the fix has to live in before_model, not in the terminal UI layer where the existing warning already sits. It's also worth being honest about what this doesn't fix: the check runs before a model call, using the cost checkpointed from the previous one. A single turn can still overshoot the cap if the cap is 5.00 and the agent is at 4.99. The next call still happens in full and might land at 6.00 before the halt fires on the turn after. That's not a bug so much as an inherent property of checking after the fact rather than metering mid-request, and it's the same tradeoff that ModelCallLimitMiddleware makes for call counts. A cap is a backstop against runaway, unattended spend. It's not a precise billing guarantee down to the last cent. Takeaways The general lesson here isn't really about cost. It's that a warning and a limit are two different features wearing the same clothing, and it's easy to ship the first while believing you've shipped the second. The warning reads the same number, uses the same word ("threshold"), and looks like it's doing the same job right up until someone isn't in the room to read it. The tell is always the same: does the check return a value the system has to act on, or does it just call something with "notify" in the name? Second, a fix that only works in-process is only half a fix once your architecture has a subprocess boundary in it. The six-hop threading here wasn't extra caution, but it was the actual scope of the problem, and skipping any one hop would have shipped a flag that silently does nothing. Third, if you can run the real thing end-to-end for two cents, there's no good reason to trust a mock's word for whether a fix actually works.

By Ninaad Rao DZone Core CORE
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do

For the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening. The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified. This is a significant architectural shift. The emerging pattern looks less like: Plain Text prompt → LLM → source code → trust it (the infamous vibe ding pattern) and increasingly like: Plain Text intent ↓ LLM ↓ restricted set of operations ↓ deterministic tools and validation ↓ result Several recent developments in mainstream developer tooling point in exactly this direction. Permission Is Becoming Separate From Intelligence On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations. Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals. That distinction is more profound than it may initially appear. The question is no longer merely: Can the agent perform this operation? It is: Is this agent authorized to perform this operation in this environment? Capability and authority are different things. A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so. A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities. Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering. The Model Is Becoming Replaceable Another development makes the separation even clearer. GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation. This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go. Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear. Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is: Plain Text stable environment stable tools stable constraints stable validation ↑ interchangeable models stable environment stable tools stable constraints stable validation ↑ interchangeable models The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system. Agents Are Separating From Editors VS Code's Agent Host architecture points in another related direction. The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts. Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code: Plain Text human ↓ editor ↓ language server ↓ compiler Agentic development introduces another participant: Plain Text human intent ↓ agent ↓ semantic tools ↓ compiler / tests / environment The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important. A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things. Verification Is Moving From Opinion to Execution A fourth development may ultimately be the most important. GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct. This should sound obvious. We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense. A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system. A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes: Plain Text generate ↓ compile ↓ test ↓ inspect diagnostics ↓ repair ↓ repeat That is substantially more robust than asking a model to inspect its own output and say whether it looks right. The role of the LLM is reasoning. The role of deterministic software remains enforcement. The Interesting Convergence These developments come from different parts of the development stack, but they point toward the same decomposition. An AI programming environment increasingly contains at least four distinct elements: A model that reasons and generatesA vocabulary of operations available to itA capability policy defining which operations it may useDeterministic mechanisms that decide whether the result is valid None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design. What If We Put the Boundary Into the Language? Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted. There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes? This is the idea I have been exploring with an open-source project called BUBAS. BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment. There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary. An order-processing application could, for example, expose operations such as: Plain Text LOAD_ORDER ORDER_TOTAL CUSTOMER_RISK APPROVE REJECT REQUEST_APPROVAL An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed. If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that. Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem. This Is Not a Sandbox Make the distinction carefully. A restricted language does not magically make its host application safe. If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong. The useful guarantee is narrower: Generated business logic can only name operations that the application deliberately made part of its vocabulary. That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment. These operate at different semantic levels. Domain Capabilities Are More Interesting Than Operating-System Capabilities Operating-system permissions are necessary, but business applications eventually need a richer vocabulary. Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful. It still does not answer questions such as: May this program approve an order?May it request approval but not approve directly?Can it read customer risk information?Can it initiate a payment?Can it calculate a premium but not change the underlying policy? Those are domain capabilities. General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction. A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations. We May Be Seeing the New Shape of the AI Programming Stack None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software. But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change. The industry is gradually replacing: Trust the model. with: Give the model precisely defined capabilities and verify what it produces. That is a much more promising engineering principle. BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.

By Peter Verhas DZone Core CORE
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.

By Akhil Madineni DZone Core CORE
Beyond Screenshots: Building Replayable Production Diagnostics for Hard-to-Reproduce Bugs
Beyond Screenshots: Building Replayable Production Diagnostics for Hard-to-Reproduce Bugs

Production bugs frequently arrive with too little evidence. A screenshot captures the final visual state, a crash report identifies a failing stack, and a support ticket describes what appeared to happen. None reliably explains the sequence that produced the failure. Modern applications are asynchronous systems driven by navigation, network responses, feature flags, background work, local persistence, and changing UI state. A useful production bug report therefore needs more than the final frame. It needs a bounded, privacy-safe execution history that reconstructs the path leading to failure. The foundation of a replayable bug report is a semantic event stream. Continuous video recording is expensive, difficult to search, and likely to capture information unrelated to diagnosis. Structured events are smaller and describe meaningful transitions directly. Navigation changes, button actions, state mutations, network outcomes, feature flag evaluations, lifecycle transitions, and persistence failures can use a common event model. Java recordEvent( "ui.action", "checkout.submit", Map.of("screen", "Checkout", "cartState", "ready") ); The event records behavior rather than pixels. Stable identifiers such as checkout.submit are preferable to coordinates or view hierarchy paths because layouts change between releases. Each event should also contain a session identifier, application version, timestamp, and monotonic sequence number. Those fields transform disconnected observations into an ordered execution history. Continuous recording immediately creates a resource constraint. Retaining every event for an entire session eventually consumes excessive memory or storage. A bounded ring buffer solves this by retaining only the most recent diagnostic window. New events replace the oldest once the configured limit is reached. Java private final Deque<ReplayEvent> events = new ArrayDeque<>(); private static final int LIMIT = 500; public synchronized void record(ReplayEvent event) { if (events.size() == LIMIT) events.removeFirst(); events.addLast(event); } Production implementations can enforce both event count and byte size limits because payload sizes vary. High-frequency signals also require control. Scroll offsets, animation callbacks, connectivity heartbeats, and repeated state notifications can overwhelm useful evidence. Sampling or coalescing these signals keeps the buffer focused on transitions that materially affect application behavior. Ordering deserves separate treatment because wall-clock timestamps alone are unreliable. Device clocks can change, concurrent callbacks can receive identical timestamps, and asynchronous tasks can finish in an order different from their creation order. Assigning an atomic sequence number when each event enters the recorder establishes deterministic local ordering. Wall-clock time remains valuable for correlation with server logs, while the sequence number determines authoritative ordering inside the client session. Network operations are especially valuable during reconstruction, but complete request and response bodies usually are not. Recording the HTTP method, normalized route, status code, duration, retry count, failure category, and correlation identifier provides useful evidence without copying sensitive payloads. Java recordNetwork( "POST", "/checkout", statusCode, elapsedMillis, correlationId ); The correlation identifier connects client replay with backend observability. When trace context propagates through an API gateway and downstream services, the same failed interaction can be followed beyond the device. OpenTelemetry context propagation supports this model through trace and span context carried across execution boundaries. A replay system becomes substantially more useful when a client event can lead directly to the corresponding backend trace instead of creating an isolated client-side observability system. Events alone may still be insufficient when identical actions behave differently under different runtime conditions. Selective state snapshots fill that gap. A snapshot should not serialize the complete application object graph. It should capture small diagnostic facts that influence behavior, such as authentication status, active feature flags, connectivity mode, pending operation counts, cache generation, lifecycle state, and other safe application state. Java recordSnapshot(Map.of( "screen", "Checkout", "network", "cellular", "paymentState", "submitting", "feature.checkoutV2", true )); Snapshots can be generated at meaningful boundaries such as screen entry, transaction start, backgrounding, synchronization completion, or error detection. During reconstruction, they explain conditions surrounding an event without attempting to reproduce every byte of runtime memory. This keeps diagnostic bundles small while preserving state likely to affect execution. Privacy must be enforced during capture rather than treated as an upload-time cleanup operation. Session replay platforms commonly provide client-side masking because sensitive values should not leave the application in their original form. The same principle applies to a custom recorder. Event attributes should follow explicit allowlists, while passwords, authorization headers, tokens, payment information, message contents, email addresses, and unrestricted text fields should be excluded by default. Java private String sanitize(String key, String value) { if (SENSITIVE_KEYS.contains(key)) return "[REDACTED]"; return value; } Key filtering alone is insufficient because sensitive values can appear under unexpected field names. Stronger implementations can combine schema allowlists, endpoint-specific policies, value-pattern detection, maximum lengths, and explicit data classifications. Redaction should occur before information enters the ring buffer. Once sensitive data has reached memory, disk, crash attachments, or telemetry queues, later sanitization becomes considerably harder to guarantee. The recorder must also survive the failure being diagnosed. An entirely in-memory history disappears during process termination, watchdog kills, native crashes, or operating system eviction. Periodic checkpointing to a small protected file allows the next launch to recover the tail of the previous session. Writes should remain asynchronous and bounded so diagnostic instrumentation does not introduce latency or instability into normal execution. Crash time behavior should remain minimal. Attempting complex serialization after a fatal condition can itself fail. A safer design periodically persists compact checkpoints during healthy execution, and treats crash handling as a final marker whenever possible. On the next launch, the previous checkpoint, crash metadata, application version, device characteristics, and correlation identifiers can be assembled into a diagnostic bundle. Upload requires failure handling as well. A diagnostic system cannot assume connectivity exists immediately after a crash or restart. Bundles can enter a small persistent queue and upload when network conditions permit. Successful delivery removes the local copy, while repeated failures follow bounded retry and retention policies. This prevents diagnostic infrastructure from becoming another source of uncontrolled storage consumption or retry storms. Replay does not need to mean pixel-perfect visual reproduction. For engineering diagnosis, a deterministic event timeline is often more useful. Plain Text 10:42:11.018 screen.enter Checkout 10:42:14.201 ui.action checkout.submit 10:42:14.233 network.start POST /checkout 10:42:15.107 network.finish status=200 10:42:15.116 persistence.error order_write_failed 10:42:15.120 ui.state payment=failed A screenshot from this incident shows only a failed checkout screen. The timeline reveals something fundamentally different, as the server operation succeeded, but local persistence failed afterward. That distinction changes the investigation immediately. Instead of examining API availability or retry behavior, diagnosis can move directly toward local storage, transaction handling, or post-response state transitions. Structured replay can extend beyond manual inspection. Development builds can consume sanitized event sequences to configure test doubles, restore relevant feature flags, reproduce network outcomes, and drive selected state transitions. Full determinism is not always possible because operating system scheduling, race conditions, third-party services, and timing-sensitive behavior introduce variability. Even without executable replay, a causal timeline drastically reduces the search space by preserving conditions that conventional logs frequently lose. The recorder should remain diagnostic infrastructure rather than becoming another analytics pipeline. Capturing every possible signal increases CPU usage, storage requirements, privacy exposure, and noise. A small event vocabulary is usually more effective, just like navigation transitions, meaningful interactions, network starts and outcomes, state changes, persistence operations, lifecycle events, and failures. Instrumentation quality matters more than event volume. A practical rollout can begin around workflows responsible for the most expensive production failures rather than instrumenting every screen. Checkout, authentication, synchronization, uploads, and other difficult flows can receive semantic events and snapshots first. When an incident demonstrates that an important transition is missing, the event model can evolve deliberately. This keeps the recorder maintainable and prevents instrumentation from becoming an uncontrolled collection of arbitrary log statements. Screenshots remain useful evidence, but they should represent one frame inside a richer diagnostic record rather than the entire debugging strategy. A bounded semantic event buffer, deterministic ordering, selective state snapshots, trace correlation, capture time redaction, crash-safe persistence, and resilient upload can transform vague production reports into reconstructable execution histories. The result is more than improved logging. It is a practical application flight recorder that preserves the context normally lost at the exact moment a difficult production failure occurs.

By Uthej Mopathi DZone Core CORE
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities

An AI agent rarely follows a long plan exactly as first proposed. Tool results expose missing facts, external systems change, policies arrive through human input, and a model may discover that an earlier assumption was wrong. ReAct-style agents were explicitly designed around this interleaving of reasoning, action, observation, and plan updates rather than one immutable plan. The engineering difficulty begins when those actions are durable side effects. In Temporal, a completed Activity is not an intention that can be edited out of a revised plan; its completion and result are part of the Workflow Event History. Replanning therefore has to reconcile a new decision with an already-real past. A Plan Is Intent; History Is Fact The cleanest mental model separates plan state from committed execution state. The plan is provisional: it represents what the agent currently intends to do. Committed state represents facts established by completed Activities, accepted external messages, and other durable events. Temporal stores Workflow history as an append-only sequence and uses that history to reconstruct Workflow state during replay. An ActivityTaskCompleted event contains the serialized Activity result, so a later model decision cannot make that completion disappear. That distinction changes the shape of the agent loop. Instead of asking an LLM to generate a complete plan and then blindly executing it, each replanning call should receive the goal, the latest observations, the prior plan revision, and a normalized set of committed effects. The model may abandon remaining steps, but completed steps should enter the next prompt as constraints on the current world state. Java public AgentResult run(Goal goal) { while (!goalSatisfied(goal, committed)) { Plan plan = planner.replan( new PlanningContext(goal, revision, committed)); for (Action action : plan.readyActions()) { if (alreadySatisfied(action, committed)) { continue; } ActionReceipt receipt = tools.execute(action, action.operationId()); committed.add(receipt); if (receipt.requiresReplan()) { break; } } revision++; } return summarize(committed); } In this shape, planner.replan(...) and tools.execute(...) are Activities, not arbitrary network calls from Workflow code. That matters because Workflow code must remain replay-safe, while non-deterministic operations such as LLM invocations belong behind durable boundaries. Temporal records Activity results in history and returns those recorded results during replay instead of re-running the external call as part of Workflow replay. Temporal’s Java AI integration applies the same principle by executing model calls and external tools through Activities, while deterministic operations may remain inside the Workflow. Replanning Should Consume Effects, Not Merely Step Status A boolean such as completed=true is usually too weak for reconciliation. A completed tool call should return an execution receipt describing the externally relevant postcondition: resource identifiers, amounts, versions, timestamps when materially relevant, and whether a compensating operation exists. The agent then reasons from effects rather than from labels such as “step three succeeded.” Consider a reservation agent that initially plans to reserve inventory, create a shipment, and notify a customer. Inventory reservation succeeds, but shipment creation reports a destination restriction. A revised plan must not schedule another reservation merely because the original plan was discarded. The durable input to replanning should state that a specific reservation already exists and identify it. The next plan might change the shipping method, release the reservation, or escalate for approval, but it should not pretend that the reservation never happened. A second guard is a planning epoch. External input can arrive while a planning Activity is outstanding; Temporal schedules Workflow Tasks when Signals or Updates arrive, and those messages can mutate durable Workflow state. A plan generated from revision 12 should therefore not be committed blindly if the state has advanced to revision 13 before the planner returns. The planner result can carry the revision used as its input, and the Workflow can discard stale output and request a fresh plan. This is optimistic concurrency control applied to agent reasoning rather than database rows. Retries make stable identity equally important. Temporal recommends idempotent Activities because Activities can be retried, and its documentation specifically notes that idempotency keys are appropriate for critical side effects. A plan revision must therefore not double as an operation identity. A reservation that remains semantically the same across replans should retain the same business operation key; a genuinely new reservation should receive a new key. Java public ReservationReceipt reserve(ReservationCommand command) { return inventory.reserve( command.sku(), command.quantity(), command.operationId()); } Here, operationId is passed through to a downstream API that enforces idempotency. The receipt returned to the Workflow becomes durable evidence of the reservation. This also prevents a subtle failure mode in which a model repeats a tool call after losing conversational context even though the Workflow already possesses proof of the earlier effect. Cancellation and Compensation Are Forward Actions A changed plan often creates pressure to “cancel the old step,” but cancellation has a precise boundary. Cancellation can stop work that is still in flight; it cannot retroactively cancel an Activity that has already completed. For long-running Activities, Temporal delivers cancellation cooperatively through Activity heartbeats, so heartbeat configuration determines how promptly an Activity can observe a cancellation request. Once an effect has committed, reconciliation becomes a domain operation. If the effect is reversible, compensation is normally the correct mechanism. Temporal documents the Saga pattern as a sequence of local operations paired with compensating actions, typically executed in reverse order when later work makes earlier effects undesirable. The important semantic point is that compensation creates new history. It does not rewrite old history. A released reservation follows a created reservation; a refund follows a charge; a revocation follows a grant. Java private void reconcile(Plan revisedPlan) { for (ActionReceipt receipt : compensationsInReverseOrder(revisedPlan, committed)) { ActionReceipt reversal = tools.compensate(receipt); committed.add(reversal); } } Not every action has a true inverse. An email cannot be unsent, an external party may already have observed a published event, and a physical operation may have crossed an irreversible boundary. Such effects should be modeled as facts that constrain future planning, not as failures of rollback. The revised plan can issue a correction, create a follow-up notification, or route the case to a human decision, but the historical effect remains part of the state presented to the agent. Durable Agents Need Forward-Only Semantics Replanning also needs to stay distinct from Workflow code versioning. A model changing its runtime plan is ordinary application behavior: a new planning Activity produces a new durable decision after new evidence arrives. Changing deployed Workflow code is different because replay must still produce commands compatible with existing Event History; Temporal provides versioning and patching mechanisms for those code changes. Mixing the two concepts leads to brittle systems in which model variability is handled as deployment variability or, worse, non-deterministic Workflow logic. External corrections fit the same forward-only model. Temporal Signals and Updates can change running Workflow state, and accepted messages become durable inputs that can trigger another planning turn. For very long-running agents, Continue-As-New creates a fresh Event History while carrying forward explicit application state, which makes the committed-effect ledger an important part of the continuation payload rather than transient model memory. The central design rule is simple: an agent may revise intentions at any time, but durable execution never revises facts. Completed Activities should be represented as committed effects with stable identities and useful receipts; in-flight work may be canceled cooperatively; reversible effects may be compensated; irreversible effects must constrain the next decision. Temporal’s history then becomes more than a recovery mechanism. It becomes the authoritative boundary between what the agent merely planned and what the surrounding world has already observed. That boundary allows adaptive AI behavior without sacrificing replay safety, idempotency, auditability, or operational correctness.

By Akhil Madineni DZone Core CORE
Building a Product Recommendation Engine With Neo4j — No ML Library Required
Building a Product Recommendation Engine With Neo4j — No ML Library Required

When many developers think about recommendation engines, they think of machine learning: collaborative filtering models, matrix factorization, embedding vectors, and training pipelines. What surprises many people is that you can build a genuinely useful recommendation system with nothing more than a graph database and several Cypher queries. No scikit-learn, no TensorFlow, no model training. Just the natural structure of the data doing the work. In this article, we'll build a product recommendation engine on top of Neo4j Aura using two Jupyter notebooks. The first generates a realistic synthetic dataset and loads it into Aura. The second runs four recommendation queries directly in Cypher and visualizes the results with Plotly. Everything runs locally in a Python virtual environment against a free cloud Neo4j instance. The full source code is available on GitHub. Why Graphs Are a Natural Fit for Recommendations The core intuition behind most recommendation approaches is relationship: this customer bought that product, those products appear together in the same order, this product shares attributes with that one. In a relational database, capturing these relationships means multiple self-joins across large tables. A query like "find products bought by customers who also bought what this customer bought" quickly becomes difficult to write and expensive to execute at scale. In a graph, that same question is a traversal. We follow edges from a customer to the products they purchased, hop across to other customers who share those products, and collect what else those customers bought. The query is short, the intent is clear, and the graph engine is optimized for exactly this kind of path-following work. Prerequisites AuraDB is Neo4j's fully managed cloud database. A free tier is available with no credit card required. Sign up at Get Started for Free.Create a new AuraDB Free instance.When the instance is created, download or note the credentials — the connection URI, username, and password.Once the instance is running, open the Query tab and connect to the instance.Confirm it's empty with MATCH (n) RETURN count(n) which should return 0 A virtual environment is highly recommended. For example: Shell python3 -m venv ~/recommendation-engine-env source ~/recommendation-engine-env/bin/activate Before starting Jupyter, export the connection details as environment variables in your shell: Shell export NEO4J_URI="neo4j+s://xxxx.databases.neo4j.io" export NEO4J_USERNAME="your_username_here" export NEO4J_PASSWORD="your_password_here" The Graph Model Before we write any code, let's define the graph. We have four node types and three relationship types. Nodes Customer – id, name, email, city, country.Product – id, name, description, price.Category – name (e.g., Electronics, Clothing, Books).Tag – name (e.g. "wireless", "eco-friendly", "premium"). Relationships (:Customer)-[:PURCHASED {order_id, quantity, order_date}]->(:Product) — order metadata lives on the relationship rather than a separate Order node, which keeps our Cypher clean.(:Product)-[:BELONGS_TO]->(:Category)(:Product)-[:TAGGED_WITH]->(:Tag) The decision to put order_id, quantity and order_date on the PURCHASED relationship is worth discussing. It means a single customer can have multiple PURCHASED relationships to the same product (each with a different order_id) and we can group by order_id to find products that appeared together in the same basket — which is exactly what our co-purchase query needs. Figure 1 illustrates exactly this point, as we have a customer, two products, and the same order_id. Figure 1. Shared order_id enables co-purchase queries Notebook 1: Data Generation and Loading Rather than sourcing an external dataset, we'll generate synthetic data using Faker. This keeps the notebook fully self-contained, and readers can run it as-is without downloading anything. We'll generate 2,000 customers, 500 products across 15 categories, and 20,000 orders. Each order is a basket of several products sharing the same order_id — this is the key design decision that makes the frequently-bought-together query work. With an average basket of 3 products, we end up with around 60,000 PURCHASED relationships in the graph. Realistic Product Names Faker's default catch_phrase() method produces output like "Proactive exuding encoding" — readable enough for a demo but not really useful in an article. Instead, we define a PRODUCT_VOCAB dictionary keyed by category, each containing lists of adjectives, nouns, use cases, and benefit statements. A product name is then a simple combination, as follows: Python def make_product_name(category): vocab = PRODUCT_VOCAB[category] adj = random.choice(vocab["adjectives"]) noun = random.choice(vocab["nouns"]) return f"{adj} {noun}" def make_product_description(category, name): vocab = PRODUCT_VOCAB[category] use_case = random.choice(vocab["use_cases"]) benefit = random.choice(vocab["benefits"]) return f"The {name} is designed for {use_case}. {benefit}." This gives us names like "Wireless Noise-Canceling Earbuds," "Organic Ground Coffee" and "Ergonomic Lumbar Support Cushion" — realistic enough to make the recommendation output meaningful. Basket-Based Order Generation Each order picks a random customer, generates a unique order_id, and samples several products into a basket. We then flatten the basket into individual order lines, each carrying the shared order_id: Python orders = [] for _ in range(NUM_ORDERS): order_id = str(uuid.uuid4()) customer = random.choice(customers) order_date = (start_date + timedelta(days=random.randint(0, 730))).strftime("%Y-%m-%d") basket = random.sample(products, k=random.randint(2, 4)) for product in basket: orders.append({ "order_id": order_id, "customer_id": customer["id"], "product_id": product["id"], "quantity": random.randint(1, 5), "order_date": order_date }) Loading Into Aura Data loading is in batches of 100 using MERGE statements. To show progress during the load, we'll use tqdm as ~60,000 order lines can take several minutes, and the progress bars make it easy to see what's happening: Python with driver.session() as session: customer_batches = range(0, len(customers), BATCH_SIZE) for i in tqdm(customer_batches, desc="Loading customers", unit="batch", colour="#1f77b4"): session.execute_write(load_customers, customers[i:i+BATCH_SIZE]) product_batches = range(0, len(products), BATCH_SIZE) for i in tqdm(product_batches, desc="Loading products ", unit="batch", colour="#1f77b4"): session.execute_write(load_products, products[i:i+BATCH_SIZE]) for product_id, tags in tqdm(product_tags.items(), desc="Loading tags ", unit="product", colour="#1f77b4"): session.execute_write(load_tags, product_id, tags) order_batches = range(0, len(orders), BATCH_SIZE) for i in tqdm(order_batches, desc="Loading orders ", unit="batch", colour="#1f77b4"): session.execute_write(load_orders, orders[i:i+BATCH_SIZE]) A verification query at the end confirms the counts. The Four Recommendation Queries Notebook 2 runs four Cypher queries against the loaded graph, each implementing a different recommendation strategy. Before running any query, we fetch a stable seed customer, product, and category: Python with driver.session() as session: customer = session.run(""" MATCH (c:Customer) RETURN c.id AS customer_id, c.name AS customer_name ORDER BY c.name ASC LIMIT 1 """).single() product = session.run(""" MATCH (p:Product)<-[r:PURCHASED]-() RETURN p.id AS product_id, p.name AS product_name, count(r) AS order_count ORDER BY order_count DESC LIMIT 1 """).single() top_cat = session.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category) RETURN cat.name AS category, count(p) AS total ORDER BY total DESC LIMIT 1 """).single() We pick the alphabetically first customer for consistency, the most-purchased product to ensure co-purchase data exists, and the category with the most products for the trending query. This makes the notebook reproducible across runs. Query 1: Collaborative Filtering The classic "customers who bought this also bought" approach. We find customers who share at least one purchased product with the seed customer, then collect what else those customers bought — excluding anything the seed customer already purchased. Python def collaborative_filtering(tx, customer_id, limit=5): result = tx.run(""" MATCH (target:Customer {id: $customer_id})-[:PURCHASED]->(p:Product) <-[:PURCHASED]-(other:Customer)-[:PURCHASED]->(rec:Product) WHERE NOT (target)-[:PURCHASED]->(rec) RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(other) AS score ORDER BY score DESC, id ASC LIMIT $limit """, customer_id=customer_id, limit=limit) return result.data() The score is the number of other customers whose purchasing overlap with our target customer also led them to buy the recommended product. A higher score means more customers in the overlap group bought it, making it a stronger signal. In Cypher, the traversal reads almost like the description: start at the target customer, follow PURCHASED edges to products, hop to other customers who bought the same products, then follow their PURCHASED edges to new products. Query 2: Frequently Bought Together This query finds products that appeared in the same order as the seed product. The key is matching on order_id across two PURCHASED relationships from the same customer: Python def frequently_bought_together(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})<-[r1:PURCHASED]-(c:Customer) -[r2:PURCHASED]->(other:Product) WHERE r1.order_id = r2.order_id AND other.id <> $product_id RETURN other.id AS id, other.name AS product, other.price AS price, count(c) AS frequency ORDER BY frequency DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The WHERE r1.order_id = r2.order_id clause is what makes this work. It constrains the traversal to only consider cases where both products were part of the same order, not just bought by the same customer at different times. frequency counts how many distinct customers placed an order containing both products together. Query 3: Content-Based Filtering Rather than looking at purchase behavior, this query finds products similar to the seed product based on shared tags. The more tags two products have in common, the more similar they are: Python def content_based(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})-[:TAGGED_WITH]->(t:Tag) <-[:TAGGED_WITH]-(rec:Product) WHERE rec.id <> $product_id RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(t) AS shared_tags ORDER BY shared_tags DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The traversal goes outward from the seed product through its tags, then back inward to any other product that shares those same tags. count(t) gives the number of shared tags, which serves as a simple but effective similarity score. This approach works without any purchase history, making it useful for recommending products to new customers or for newly listed products with no order data yet. Query 4: Trending in Category This query finds the most purchased products in the top category within a fixed date window. In our case, this is from 2024-10-01 onwards: Python def trending_in_category(tx, category_name, cutoff="2024-10-01", limit=5): result = tx.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category {name: $category_name}) MATCH (:Customer)-[r:PURCHASED]->(p) WHERE date(r.order_date) >= date($cutoff) RETURN p.id AS id, p.name AS product, p.price AS price, count(r) AS purchases ORDER BY purchases DESC, id ASC LIMIT $limit """, category_name=category_name, cutoff=cutoff, limit=limit) return result.data() We use date() conversion on the stored string order_date to enable date comparison. count(r) counts individual PURCHASED relationships rather than distinct customers, so a customer who bought the same product multiple times within the window is counted each time — reflecting genuine demand volume rather than unique buyer count. Notebook 2: Results Each query outputs a table followed by a Plotly horizontal bar chart. Here are the results for our seed data. Collaborative Filtering Figure 2 returns five products. The top recommendation is An Introduction to Public Speaking, driven by the number of customers whose purchasing overlap with Aaron Boyd also led them to buy it. Heavy-Duty Cable Management Box and Educational Coding Robot follow closely, showing that the overlap group bought broadly across categories rather than clustering in one area. Figure 2. Collaborative filtering Frequently Bought Together Figure 3 shows products co-purchased with the Durable Grooming Brush in the same order basket. The top results — Waterproof Hammock and Natural Body Lotion at frequency 4, followed by Adjustable Lumbar Support Cushion, Sugar-Free Collagen Powder and Slim-Fit Hiking Vest at frequency 3 — show which products most commonly appeared alongside the seed product in the same order. The cross-category spread here (Beauty, Outdoor, Clothing, Health, Office) is a feature of random synthetic data; in a real system, we'd expect more category clustering. Figure 3. Frequently bought together Content-Based Filtering Figure 4 finds products sharing the most tags with the seed product. All five results share 2 tags with the Durable Grooming Brush — Smart Mechanical Keyboard, Waterproof Toiletry Bag, Ergonomic Whiteboard, Cold-Pressed Hot Sauce, and Durable Dumbbell Pair. The cross-category reach (Sports, Food & Drink, Office, Travel, Electronics) illustrates the tag graph doing its job: shared attributes like "durable" or "waterproof" create similarity links that cross category boundaries, which is useful for surface-level discovery recommendations. Figure 4. Content-based filtering Trending in Category Figure 5 shows the top 5 products in Toys — the category with the most products in our graph — with purchase counts from 2024-10-01 onwards. Battery-Free Coding Robot leads, followed by Battery-Free Building Blocks Set, Interactive Remote Control Car, Wooden Magnetic Drawing Board, and Creative Puzzle Game. The scores are tight here, which makes sense because within a single category over a fixed time window, popular products tend to cluster around similar purchase volumes. Figure 5. Trending Summary We've built a working product recommendation engine using nothing but Neo4j, Cypher, and a few Python libraries. No ML framework, no training data, no model deployment. The four queries cover the most common recommendation patterns in production systems: Collaborative filteringCo-purchase analysisContent similarityTrending detection The graph model is the foundation that makes this possible. Storing orders as relationships with properties means co-purchase queries are a natural traversal rather than a complex join. Adding tags as nodes means similarity queries are just path-matching. Because everything lives in the same graph, we can also combine these approaches. For example, filtering collaborative filtering results by tag similarity using a single extended Cypher query. The full source code is available on GitHub.

By Akmal Chaudhri DZone Core CORE
Why Incident Response Needs Memory, Not Just Intelligence
Why Incident Response Needs Memory, Not Just Intelligence

Every production incident starts with a simple question: "Has this happened before?" I've lost count of how many incident bridges I've joined where that question came up within the first few minutes. Before anyone proposes restarting a service or rolling back a deployment, someone inevitably starts searching. They look through Slack conversations from previous outages, browse old postmortems, compare dashboards with similar incidents, or dig through runbooks to see whether another team has already solved the same problem. Do you notice what's happening here? The engineers aren't trying to demonstrate how much they know about distributed systems. They're trying to remember. That observation has become increasingly important as AI assistants find their way into engineering organizations. Today's large language models are remarkably good at explaining Kubernetes concepts, debugging stack traces, writing SQL queries, or summarizing log files. Those capabilities are valuable, but they only solve part of the problem. During an incident, reasoning is rarely the bottleneck. Finding the right context is. The engineer who resolves an outage the fastest isn't always the one with the deepest theoretical knowledge. More often, it's the engineer who remembers that a similar issue occurred eight months ago after a database failover, or who knows that a particular service has historically exhibited the same failure pattern after specific deployment changes. Experience is a form of memory. If we want AI to become a trusted operational partner instead of just another chatbot, we need to think about memory as carefully as we think about intelligence. Intelligence Answers Questions. Memory Solves Problems. Large language models are exceptional at answering questions because they've been trained on an enormous amount of public knowledge. Ask an LLM to explain consensus algorithms, Kubernetes scheduling, or distributed tracing, and you'll probably receive a detailed, technically accurate explanation within seconds. Production incidents, however, ask very different questions. Instead of asking: What is Kubernetes? Engineers ask: Why did our Kubernetes cluster start failing after yesterday's deployment?Has this service failed in the same way before?Which runbook actually worked the last time?Who owns this dependency today?What changed in the last hour that could explain this behavior? Those aren't questions about computer science. They're questions about organizational memory. The answers don't exist in a foundation model's training data because they're unique to every organization. They live in deployment histories, incident timelines, internal documentation, architecture decisions, monitoring dashboards, chat conversations, and postmortems accumulated over years of operating software. An AI that understands distributed systems but lacks access to this operational history is like an experienced consultant joining your incident bridge for the first time. It may offer useful suggestions, but it doesn't know your environment, your systems, or your team's accumulated experience. That's why intelligence alone isn't enough. Every Incident Is a Search Problem One pattern I've noticed is that incident response often looks less like debugging and more like information retrieval. Think about what engineers actually do during the first fifteen minutes of a major incident. One person opens dashboards to identify where the failure started. Another compares the current deployment with the previous version. Someone searches Slack for keywords that resemble the current symptoms. Another engineer opens the last postmortem for the affected service. Meanwhile, the incident commander tries to understand which teams need to be involved. None of these activities involve writing complex algorithms. They're all attempts to reconstruct context. If you mapped the engineer's workflow, it might look something like this: Alert → Metrics → Logs → Deployment History → Previous Incidents → Runbooks → Slack Discussions → Architecture Documentation → Decision The common thread is that engineers are constantly retrieving information before making decisions. That retrieval process is exactly where AI can provide the most value—not by replacing engineering judgment, but by dramatically reducing the time required to gather relevant context. Not All Memory Is the Same When we talk about memory in AI, it's easy to think only about conversation history or a vector database. In practice, incident response depends on several different kinds of memory, each answering a different set of questions. Incident memory This is the collective history of operational failures. Previous incidents, timelines, root causes, postmortems, and lessons learned all fall into this category. During an outage, one of the first questions engineers ask is whether they've seen the problem before. An AI that can retrieve similar incidents and explain how they were resolved immediately provides value because it shortens the investigation. Operational memory Runbooks, playbooks, escalation procedures, and service ownership represent another form of memory. These artifacts capture how an organization expects engineers to respond under different circumstances. Instead of generating a generic remediation plan, an AI can recommend the procedure that has already been validated by the organization. Infrastructure memory Production systems change constantly. Deployments, feature flags, infrastructure updates, configuration changes, and dependency upgrades all influence system behavior. Understanding what changed recently is often more useful than understanding how a technology works in theory. Organizational memory Some of the most valuable operational knowledge never reaches formal documentation. Engineers discuss recurring issues in Slack, record architectural decisions in design documents, and exchange troubleshooting tips during retrospectives. Over time, this becomes institutional knowledge that experienced engineers rely on instinctively. AI should be able to surface that knowledge instead of forcing every engineer to rediscover it. Memory Changes the Quality of Recommendations Imagine two AI assistants responding to the same latency alert. The first assistant says: CPU utilization is high. Consider restarting the service. It's not necessarily wrong, but it's also not particularly helpful. Now imagine a second assistant with access to organizational memory says: A similar incident occurred three months ago after deployment version 6.4. During that incident, restarting the service temporarily reduced latency, but the underlying cause was an inefficient database query introduced by the deployment. The query was reverted, and latency returned to normal within six minutes. A deployment with similar changes occurred eighteen minutes before the current alert. I recommend validating query performance before restarting the service. Neither assistant is more intelligent in the traditional sense. The second assistant is simply making better use of memory. That additional context changes the recommendation from a generic suggestion into operational guidance grounded in the organization's own experience. Building Memory Into AI Systems Memory isn't a single database or a single technology. It's an architectural capability that combines multiple sources of operational knowledge into a coherent context for reasoning. A production-ready incident assistant might continuously ingest information from observability platforms, deployment pipelines, service catalogs, incident management systems, internal documentation, and communication channels. Rather than asking engineers to manually gather information from each source, the AI assembles the relevant context before generating a recommendation. The language model is still responsible for reasoning, summarization, and communication. The memory layer ensures that reasoning is grounded in facts that are specific to the organization rather than generic patterns learned during training. In many ways, this mirrors how experienced engineers work. They don't solve incidents by relying only on theoretical knowledge. They combine technical understanding with years of accumulated operational experience. That's exactly the capability our AI systems should emulate. Final Thoughts As large language models continue to improve, it's tempting to believe that more intelligence alone will solve the challenges of operational AI. My experience suggests otherwise. The most effective incident response systems aren't necessarily the ones with the largest models or the most sophisticated prompts. They're the ones that help engineers remember. They surface the right runbook, identify the last time a service failed in the same way, highlight the deployment that introduced the problem, and connect today's symptoms with yesterday's lessons. In other words, they make organizational experience accessible when it's needed most. Incident response has always been a combination of reasoning and memory. AI has made remarkable progress on the first half of that equation. The next step isn't simply building smarter models—it's building systems that remember.

By Akshay Pratinav

The Latest Data Engineering Topics

article thumbnail
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
Learn about graph databases by building an F1 teammate network from real Formula 1 data and using Cypher to connect Max Verstappen to Juan Manuel Fangio.
September 30, 2026
by Jeremy Morgan
· 205 Views
article thumbnail
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
Learn how to classify workloads, choose the right migration path, and avoid the traps that turn 6-week projects into 6-month ones.
September 30, 2026
by Jerzy Kopaczewski
· 206 Views
article thumbnail
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
Databricks and Snowflake speak the Kafka protocol, but Kafka for lakehouse ingestion is not Kafka as an event-driven architecture.
September 30, 2026
by Kai Wähner DZone Core CORE
· 435 Views
article thumbnail
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code
AI-generated code needs verifiable provenance linking intent, context, models, edits, approvals, commits, and artifacts across the software lifecycle.
September 30, 2026
by Uthej Mopathi DZone Core CORE
· 240 Views
article thumbnail
Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival
Meta launches Enterprise Platform, bringing Muse, Business Agent and AI tools to companies as it takes on Microsoft and Salesforce.
September 30, 2026
by Ai Cerrudo
· 323 Views
article thumbnail
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
Build self-healing SQL pipelines where LLMs propose repairs while deterministic validation, guardrails, and execution controls protect production systems.
September 30, 2026
by Uthej Mopathi DZone Core CORE
· 270 Views · 1 Like
article thumbnail
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
OpenAI may be testing an always-on ChatGPT assistant called “o,” with leaked references pointing to possible email functionality and ChatGPT Pro placement.
September 29, 2026
by DZone Staff
· 1,274 Views · 1 Like
article thumbnail
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A hands-on guide to Microsoft Foundry Document Intelligence SDK for extracting text, structured fields, and document data from PDFs for RAG and AI pipelines.
September 29, 2026
by Jubin Soni, FBCS DZone Core CORE
· 673 Views
article thumbnail
How to Test POST API Requests With Playwright TypeScript
Learn how to test POST API requests in Playwright with TypeScript using static JSON objects and arrays, JSON.stringify(), JSON files, and the Faker library.
September 29, 2026
by Faisal Khatri DZone Core CORE
· 484 Views
article thumbnail
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
In this article, you will learn why lines of code generated are almost meaningless as a measure of AI-assisted development.
September 29, 2026
by Gaurav Gaur DZone Core CORE
· 702 Views · 1 Like
article thumbnail
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch gives enterprise apps a standard model for long-running data processing with jobs, steps, readers, processors, writers, checkpoints, and tunable execution.
September 29, 2026
by Otavio Santana DZone Core CORE
· 597 Views
article thumbnail
The Warning That Never Stops the Agent
Deepagents tracks real session cost but only shows a toast. Nothing stops the loop from spending past a limit. A before_model hook that can jump to "end" closes that gap.
September 29, 2026
by Ninaad Rao DZone Core CORE
· 453 Views
article thumbnail
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI coding is shifting from trusting models to constraining them with permissions, tools, and deterministic checks. BUBAS applies the same idea to business logic.
September 28, 2026
by Peter Verhas DZone Core CORE
· 599 Views
article thumbnail
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
Progressive disclosure replaces giant prompts with lightweight skill summaries and on-demand instructions, keeping AI agents focused and extensible.
September 28, 2026
by Akhil Madineni DZone Core CORE
· 679 Views · 1 Like
article thumbnail
Beyond Screenshots: Building Replayable Production Diagnostics for Hard-to-Reproduce Bugs
Capture privacy-safe production event timelines to reconstruct failures, correlate backend activity, and diagnose bugs that screenshots cannot explain reliably.
September 28, 2026
by Uthej Mopathi DZone Core CORE
· 428 Views · 1 Like
article thumbnail
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities
Reconcile AI plan changes with completed Temporal Activities using versioned plans, idempotent execution, and compensation rather than retroactive rollback.
September 28, 2026
by Akhil Madineni DZone Core CORE
· 628 Views · 2 Likes
article thumbnail
Building a Product Recommendation Engine With Neo4j — No ML Library Required
A graph models customers, products, categories and tags, making collaborative filtering, co-purchase analysis, content similarity, and trending queries graph traversals.
September 28, 2026
by Akmal Chaudhri DZone Core CORE
· 597 Views
article thumbnail
Why Incident Response Needs Memory, Not Just Intelligence
LLMs are excellent at reasoning, but effective incident response depends just as much on remembering previous incidents and organizational operational context.
September 28, 2026
by Akshay Pratinav
· 592 Views
article thumbnail
Prompt Caching Doesn't Save Money on Turn One
Cache writes cost 1.25x the input price, and reads cost 0.1x. So a fresh conversation's first turn is more expensive with caching on.
September 25, 2026
by Ninaad Rao DZone Core CORE
· 1,081 Views · 1 Like
article thumbnail
Software Quality Habits and AI
Software quality depends on daily habits. Learn how small habits compound and how AI can strengthen quality while weakening engineering judgment.
September 25, 2026
by Stelios Manioudakis DZone Core CORE
· 2,266 Views · 2 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×