DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

AI/ML

Artificial intelligence (AI) and machine learning (ML) are two fields that work together to create computer systems capable of perception, recognition, decision-making, and translation. Separately, AI is the ability for a computer system to mimic human intelligence through math and logic, and ML builds off AI by developing methods that "learn" through experience and do not require instruction. In the AI/ML Zone, you'll find resources ranging from tutorials to use cases that will help you navigate this rapidly growing field.

icon
Latest Premium Content
Trend Report
Generative AI
Generative AI
Refcard #403
Shipping Production-Grade AI Agents
Shipping Production-Grade AI Agents
Refcard #401
Getting Started With Agentic AI
Getting Started With Agentic AI

DZone's Featured AI/ML Resources

Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival

By Ai Cerrudo
Meta has spent years helping businesses advertise on Facebook and Instagram and talk to customers through WhatsApp. Now it wants its AI working more deeply within those businesses. The company launched Meta Enterprise Platform on Monday, a new business focused on turning its AI models, agents, coding tools, and infrastructure into products companies can deploy themselves. CEO Mark Zuckerberg called the platform the "next major pillar" of Meta's business. Its initial technology stack includes the Muse agent, Meta Business Agent, Muse API, and Muse Code. The move puts Meta more directly into an enterprise AI market already crowded with Microsoft, Salesforce, Google, ServiceNow, and other vendors racing to make AI agents part of everyday business operations. Meta is turning its AI stack into a business platform Meta says it already helps hundreds of millions of businesses reach customers. Enterprise Platform is an attempt to expand that relationship beyond advertising and messaging. Rather than introducing a single standalone enterprise application, Meta plans to bring together several components of its AI portfolio. Muse is Meta's general-purpose AI agent, capable of taking actions such as sending emails, booking travel, and completing other multistep tasks. TechRepublic previously reported that Muse launched in the U.S. this month across iOS, Android, the web, and WhatsApp. Meta Business Agent is aimed more directly at companies. Meta says more than one million businesses already use a Business Agent on WhatsApp and Messenger to respond to customers. The agent can answer questions, recommend products, book appointments, qualify leads, and close sales. Meta also says more than one billion active threads with businesses take place across WhatsApp, Messenger, and Instagram each day. For larger organizations, Meta Business Agent Platform can connect with hundreds of systems, including Shopify, Zendesk, and Shopee. Meta says the platform includes enterprise controls, guardrails, and measurement tools that let companies define how their agents operate. Meta has also outlined broader ambitions for Business Agent. The company says it eventually wants the technology to help with tasks such as market research, product insights, calendar management, and competitive intelligence. Enterprise Platform now gives those efforts a dedicated business operation. Meta hires an enterprise software veteran to lead the push Meta recruited former MongoDB CEO Chirantan "CJ" Desai as chief enterprise platform officer. He will report directly to Zuckerberg. Before leading MongoDB, Desai oversaw product and engineering at Cloudflare and spent nearly eight years at ServiceNow, including as president and chief operating officer. "Meta Enterprise Platform will focus on turning its AI stack into products and services that companies can deploy for their own businesses," Desai said in Meta's announcement. The hire gives Meta an executive with experience building and selling software to large organizations as it moves beyond its traditional advertising and consumer-platform businesses. Microsoft and Salesforce have another AI competitor Meta is entering a market where major enterprise software companies have already spent years embedding AI into existing business platforms. Microsoft offers AI agents across Microsoft 365 Copilot and Dynamics 365, including agents that can access organizational data and take actions such as sending emails or updating records. Salesforce has similarly expanded Agentforce across its CRM products, positioning autonomous agents around sales, service, marketing, and other customer workflows. Meta starts from a different position. It does not have Microsoft's productivity suite or Salesforce's CRM footprint, but it does have massive consumer reach, business messaging platforms, AI models, agents, and large-scale infrastructure. That could give Meta an opening through customer-facing operations before it pushes further into internal enterprise workflows. The move could also provide another way for Meta to generate returns from its rapidly expanding AI infrastructure. According to Meta, the company currently expects 2026 capital expenditures of between $130 billion and $145 billion, with data center capacity among the drivers of that spending. Zuckerberg has previously said Meta sees a broader enterprise opportunity that could include APIs, business agents, services for large customers, and potentially selling compute directly. Meta still has big enterprise questions to answer For all the ambition behind the announcement, Meta Enterprise Platform is still more of an enterprise strategy than a clearly packaged software suite. Meta has not disclosed overall platform pricing, named launch customers, or explained exactly how Muse, Business Agent, Muse API, and Muse Code will be packaged and administered together. Those details will matter for IT leaders deciding whether Meta belongs alongside established enterprise vendors. Organizations will need to understand how administrators control agent access, which actions require human approval, how activity is logged and audited, and how Meta's products integrate with existing identity, security, and business systems. Meta says security and privacy will be built into its enterprise products from the outset. Its existing Business Agent Platform already includes enterprise controls and guardrails, but Meta has not yet provided the same level of detail for the broader Enterprise Platform. That leaves the biggest question unanswered: whether Meta can turn its enormous consumer reach and AI investment into an enterprise platform companies trust with core business operations. For now, Meta is no longer content with helping businesses reach customers through its apps. It wants its AI agents doing some of the work behind them, too. Editor’s note: This article originally appeared on our sister publication, TechRepublic. More
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code

Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code

By Uthej Mopathi DZone Core CORE
AI-generated software needs provenance that survives beyond the chat window. A code review can show what changed, but it rarely shows which model produced a fragment, what prompt and repository context influenced it, which agent or tool executed the change, which intent was being implemented, or whether the recorded history was altered later. Provenance fills that gap by treating generation as a supply-chain event rather than an ephemeral interaction. The core idea is already established in adjacent standards where W3C PROV models provenance through entities, activities, and agents, while SLSA records how software artifacts were produced so downstream consumers can verify expected processes and inputs. Generation Provenance as Engineering Metadata The first design rule is to separate authorship from provenance. Provenance answers where code came from and how it was produced, and it does not, by itself, determine legal ownership. The U.S. Copyright Office states that generative-AI output is copyrightable only where sufficient human-authored expressive elements exist, and that prompting alone is not enough. Employment agreements, contributor agreements, licenses, and jurisdictional law still govern ownership questions. Provenance instead supplies evidence for attribution, review, audit, and accountability. A minimal record should bind the generated artifact to the model provider, model identifier and revision, agent identity and version, prompt digest, context digests, execution trace, repository commit, intent digest, timestamp, and approving human or service identity. Hosted model aliases can change over time, so a provider-returned model identifier or immutable deployment revision is preferable to a friendly model name alone. The record should also contain cryptographic digests for generated files so later edits cannot silently inherit stale provenance. JSON { "artifact": "src/billing/CancelService.java", "sha256": "7e91...c42a", "commit": "9f3c1ad", "model": {"provider": "acme-ai", "id": "code-model", "revision": "2026-08-14"}, "agent": {"id": "repo-agent", "version": "3.7.2"}, "prompt": "sha256:18ab...90ef", "context": ["git:9f3c1ad^", "sha256:44c2...bb10"], "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1...", "approvedBy": "team:payments-reviewers" } Git commit trailers provide a low-friction place to attach pointers because Git supports structured token-value trailers at the end of commit messages. The commit should store references and digests rather than sensitive prompts themselves. Plain Text AI-Provenance: sha256:5df0...a992 AI-Model: acme-ai/code-model@2026-08-14 AI-Agent: [email protected] Prompt-Digest: sha256:18ab...90ef Context-Digest: sha256:44c2...bb10 AI-Trace: urn:uuid:2cf1... Intent-Digest: sha256:a771...0d61 File-level attribution can use a compact pointer rather than duplicating the full record. A generated region can carry a comment such as // ai-provenance: urn:gen:2cf1..., while the referenced sidecar record maps that generation to file hashes and, when needed, line ranges. This keeps source readable and prevents model metadata from becoming scattered, inconsistent comments. From SBOM to Generation BOM AI-generated code needs an additional description of the generation process. CycloneDX already supports source code, machine-learning models, component provenance, formulation describing how objects were created, and citations that attribute supplied information to entities or processes. Its ML-BOM capability records models, datasets, configurations, and provenance, while the earlier model-card framework established the broader practice of documenting model identity, intended use, evaluation, and limitations. A generation BOM can therefore be implemented as a small signed sidecar linked to the repository and release artifact, rather than inventing a second source-control system. JSON { "bomFormat": "GenerationBOM", "specVersion": "0.1", "subject": {"path": "src/billing/CancelService.java", "sha256": "7e91...c42a"}, "generator": {"agent": "[email protected]", "model": "acme-ai/code-model@2026-08-14"}, "inputs": {"prompt": "sha256:18ab...90ef", "context": ["sha256:44c2...bb10"]}, "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1..." } The intent digest is especially important. A prompt records instructions presented to a model, but an intent contract records the behavior that must remain true after generation. Such a contract can contain permitted change scope, protected behaviors, security constraints, and acceptance criteria. Provenance then connects the produced code not merely to an AI request, but to a reviewable engineering objective. This mirrors data-lineage systems such as OpenLineage, which associate runs, jobs, datasets, and extensible facets so downstream analysis can reconstruct how an output was produced. Artifact hashes alone cannot establish reproducibility when generation depends on mutable infrastructure. Provenance should therefore bind execution parameters such as model configuration, decoding settings, tool versions, retrieval indexes, and policy revisions. Capturing these values converts provenance from a historical label into a verifiable reconstruction boundary for later audits and incident analysis. Tamper Evidence and CI Enforcement Metadata becomes trustworthy only when alteration is detectable. SLSA explicitly treats provenance authenticity and digital-signature verification as mechanisms for detecting tampering, and recommends approaches that improve compromise detection, including transparency logs. Sigstore provides signing with short-lived identity-bound certificates and records signing events in Rekor, an append-only transparency log. A provenance file can be signed as a blob during CI: Shell cosign sign-blob \ --bundle generation-provenance.sigstore.json \ generation-provenance.json Verification should occur before merge or release, not after an incident. SLSA similarly emphasizes that provenance has little value unless a consumer verifies it against expected properties. Shell provctl verify \ --commit "$GIT_COMMIT" \ --require-model \ --require-context-digest \ --require-intent \ --require-signature \ --max-unattributed-lines 0 A practical gate should reject AI-marked changes when the artifact digest no longer matches, required model or agent fields are absent, the provenance signature fails, the intent contract is missing, or the trace cannot be resolved. Human-edited code should not be forced into artificial AI attribution; instead, the policy should distinguish generated, transformed, and manually authored regions. Execution traces can preserve tool calls, retrievals, test runs, and agent steps and are designed to make software-supply-chain steps transparent by recording what happened, by whom, and in what order, and their runtime-trace predicate can describe system events associated with a supply-chain step. Runtime verification closes another gap. Provenance can prove which generation path produced a deployment, but not that the resulting behavior remains correct under production conditions. Release telemetry should therefore link runtime incidents back to commit, provenance record, model revision, and intent contract. That correlation turns an AI-related defect from an unstructured forensic exercise into a query over lineage. Accountability Without Capturing Everything Capturing every prompt verbatim is usually the wrong default. Prompts and retrieved context may contain credentials, personal data, proprietary code, customer information, or licensed material. A safer design stores encrypted source material in an access-controlled evidence store and places digests, object references, retention class, and classification labels in Git-visible provenance. High-sensitivity environments can retain only keyed digests and approved summaries where reproduction is less important than proof of correspondence. Storage and performance costs also require boundaries. Full agent traces can be large, while line-level metadata can become noisy after refactoring. The durable unit should normally be a generation event bound to artifact digests and commits, with finer-grained ranges reserved for high-risk code. Developer ergonomics matter equally, as provenance capture should be automatic in IDE agents, repository bots, and CI runners rather than dependent on manual form filling. Regulation strengthens the case for disciplined records without creating a universal rule that every AI-generated source line must carry a label. The EU AI Act requires general-purpose AI model providers to maintain technical documentation and copyright-compliance policies, while NIST SP 800-218A extends secure development practices specifically for generative AI across the software lifecycle. These frameworks reinforce documentation, traceability, and governance, but a code-provenance system should be treated as engineering evidence rather than a substitute for legal analysis. AI-generated code should enter a repository with the same expectations applied to any other supply-chain artifact: origin, inputs, process, identity, integrity, and approval must be recoverable later. The strongest implementation is not a comment saying that AI was used, but a signed provenance chain linking model and agent identity, prompt and context digests, execution trace, intent contract, commit, generated artifact, review decision, and runtime evidence. Teams adopting AI-assisted development should make that chain automatic, verify it in CI, protect sensitive evidence separately, and fail closed for unattributed high-risk changes. That converts provenance from documentation into an enforceable engineering control and makes accountability possible long after the generation session has disappeared. More
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
By Jubin Soni, FBCS DZone Core CORE
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
By DZone Staff
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
By Uthej Mopathi DZone Core CORE
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1

A few months ago, one of our teams celebrated a milestone in their quarterly review. AI adoption was up. The productivity dashboard showed that developers were generating, on average, 40% more code per sprint. The tech lead showed this as a major win. Three weeks later, I was on a call for a production issue. The challenge was that errors were coming in three different formats depending on which endpoint you hit. The alerting was blind to a category of failure it had always caught before. We had a centralized exception handler. It logged context, mapped it to the right HTTP status, and pushed these alerts to our observability stack. When we investigated, we found that recent AI-assisted PRs had started introducing their own try-catch blocks inline. Each one caught exceptions locally, logged in a slightly different format, and returned a slightly different error shape. Some swallowed the exception instead of letting it propagate up to the handler that would have alerted us. Each one of those PRs was correct. Every one passed review, including the reviews I did myself. We were all checking for correctness, and consistency isn't the kind of thing that shows up in a git diff. Cleaning it up took most of a sprint. And that productivity dashboard? It counted the original generation and the cleanup as output. Twice the code, twice the "productivity," for a net loss of engineering time. The dashboard was going up while the system got worse underneath it. That's the mistake I keep seeing where teams confuse code generation with engineering progress. The LOC Trap, Reloaded Fred Brooks called this out in The Mythical Man-Month decades ago. He explicitly mentioned that measuring programming productivity by lines of code is nonsensical. Everyone agreed and then somehow forgot. So why did we rebuild the exact same dashboard the moment AI arrived? We are using the same flawed metric, with AI branding, and presenting it to boards. When a team lead reports that AI tools helped produce 40% more code, the follow-up I want to hear is “Did we actually need 40% more code?” Usually the answer is no. What we needed was the same outcomes with less effort, and effort in software lives overwhelmingly outside the act of typing. I’ll come back to that. There is a difference this time, and it is worth naming. Teams aren’t defending line counts out loud anymore. The dashboards have moved on to merged pull requests, agent tasks completed, and suggestions accepted. It is the same instinct in a unit that sounds more respectable in front of a board, counting the artifacts of work and reporting the count as progress. Why Senior Engineers Delete Code Here’s a pattern you’ll recognize if you have led engineering teams for any length of time. Your best engineers often produce fewer lines of code than anyone else on the team. Some of their most impactful weeks come out to negative line counts. That’s expertise rather than laziness. I haven’t fully worked out why the instinct for deletion over addition takes years to develop, but it does. GitClear has been measuring this rather than speculating about it. Their 2026 analysis covers 623 million changes from 2023 to 2026, and the figure that stopped me had nothing to do with volume. Moved code, their proxy for refactoring, dropped from 21% of all changes in 2022 to 3.8% by the middle of this year. Duplicated blocks are up 81% across the same window. Cross-file function calls, which is roughly what reuse looks like in a diff, are down 35%. Those three together describe a codebase that has stopped being rearranged when work goes in, and very little gets moved or deleted. When someone takes 2,000 lines of tangled logic and replaces it with 200 clean ones, that looks like a loss on any volume metric. It is an enormous win for the system, and it is exactly the activity that has gone quiet. One correction I owe, since I quoted the earlier version of this research at people for the better part of a year. GitClear’s 2024 report predicted two-week code churn would double in the AI era. It didn’t double. It went up 15%. The headline projection was too aggressive, and the part almost nobody quoted — refactoring falling off a cliff — turned out worse than predicted. Senior engineers get this intuitively. Every line of code is a liability, because every line has to be read, understood, tested, and maintained. So the best solution often makes code disappear, and that takes different forms like a well-chosen abstraction that kills duplication or a config change that removes a custom implementation. Sometimes it’s just a conversation with the PO that drops the requirement entirely. Now think about what AI coding metrics would say about this. An engineer spends a day understanding a system, realizes three services can collapse into one, and deletes 4,000 lines. By every AI productivity metric in use today, that engineer had a terrible day. In reality, they may have saved the organization months of future pain. Gergely Orosz tells a revealing story about what happens when you optimize for the wrong signal. When Uber introduced diff-count metrics, engineers started creating more, smaller changes to look productive. They flooded CI systems, driving up costs. The metric improved and engineering got worse. We are setting ourselves up for the same trap with AI-generated LOC. AI’s Tendency Toward Verbose Implementations This gets worse when you look at what AI coding tools actually excel at, which is producing plausible code quickly. That skill carries a built-in bias toward verbosity. Sometimes more code is genuinely the right call. Explicit beats implicit. A verbose but readable implementation can be better than a clever one-liner that nobody understands at 3 am when production is on fire. I’m not arguing for code golf. But AI-generated verbosity is a specific kind of bad, because it is default verbosity from ignorance of context rather than chosen verbosity for clarity. This distinction matters a lot more than I initially thought. Ask an AI assistant to implement a feature, and you’ll get a complete, working solution, longer than what an experienced developer would write, because it optimizes for correctness and completeness in isolation. It may not pick up the utility you wrote last month to do exactly this. It doesn’t realize the framework provides a one-liner if you structure the problem slightly differently. It can’t tell the difference between “I should be explicit here for readability” and “I am reinventing something that already exists three directories over.” The handler drift I opened with is the cleanest example I have of it. The AI did exactly what it was asked, every single time. Each PR added code, each one passed review on its own terms, and nothing was wrong inside any of them. What we lost lived across them. One handler gave us one error shape, and one error shape gave our alerting something to fire on. No single diff broke that, and all of them together did. GitClear tracks error-masking constructs, which is the failure mode buried in there, and they are up 47% since 2023. Our inline handlers were exactly that. I’d like to think we were an unlucky outlier, and the data says we were ordinary. I’m still not sure how you review for a property that isn’t visible in the file in front of you. Complexity as the Hidden Cost Code volume isn’t a perfect proxy for system complexity, and I acknowledged that above. But it is a directional one, and in aggregate it holds. When your codebase grows by 30-40% in a quarter without a corresponding growth in functionality, complexity is almost certainly growing with it. And system complexity is the single biggest thing determining how fast your team can move over time. I haven’t found a way around that in eighteen years. Every line of code carries ongoing costs that nobody puts on a dashboard: Cognitive load for anyone working nearbyTest coverage, without which it becomes a ticking riskReview time on every future change to itDependencies it drags in that need constant updatingMigration effort during every platform change When AI tools grow your code volume by 30-40%, each of these costs grows. The productivity gain at the moment of writing is real, and I’m not denying that. But it can be entirely eaten up by the downstream cost of maintaining a bigger, more complex system. Then there is the study I keep coming back to, and the update to it that I nearly missed. In mid-2025, METR found that experienced developers using AI coding tools took 19% longer on real-world tasks. The setup was specific. It had 16 open-source contributors working in their own repositories, code they’d lived in for years, 246 real issues averaging about two hours each, with Cursor Pro and Claude. In February 2026, METR published an update that takes a fair amount of it back. They are redesigning the experiment, and the reasons aren’t flattering to the original. Developers who would no longer work without AI declined to take part at all. Somewhere between 30% and 50% of participants avoided submitting exactly the tasks where they expected to want AI. The pay rate for the follow-up work dropped from $150 an hour to $50, which made recruitment worse again. Their own summary is that the newer data amounts to very weak evidence in either direction, and that developers are probably more sped up now, in early 2026, than the early-2025 estimate suggested. So the 19% was never a fact about AI-assisted development. It was a measurement of sixteen people in one setting, and the people who ran it now think it read low. What survives is the part that makes me think harder. Those developers believed they were 20% faster. Whatever the true effect was, it wasn’t the effect they perceived, and not one of them could feel the gap while it was happening. Reviewing and integrating suggestions consumed time none of them accounted for. Selection bias moved the headline number around, but it does not explain away a room full of experienced engineers being wrong about their own week. I could never square that finding with my own experience, because I do feel faster on certain tasks. The update moves me off the fence, slightly. Maybe I wasn’t fooling myself. I would still put no weight at all on my own estimate of how much faster I am, and that’s close to the only thing here I am confident about. None of that is an argument against the tools, only against measuring them by how much code they produce. The Strongest Number Against Me If you want to argue the other side, the best evidence available today is Microsoft’s. Early in 2026, they rolled Claude Code and GitHub Copilot CLI out across the organization and studied what happened. Tens of thousands of engineers, four months, and the ones who adopted merged roughly 24% more pull requests than the counterfactual said they would have. That is not a lab, and it is not a small group of volunteers. It’s the largest measurement of agentic coding tools anyone has published, the effect is large, and it points the right way. I take it seriously. I also notice what the unit is. The authors get there ahead of any critic. Their paper says a merged PR isn’t the same as the value it delivers, which is the argument of this entire post, conceded inside the study that is supposed to answer it. Twenty-four percent more merged pull requests is consistent with 24% more delivered value. It is equally consistent with the same work arriving in smaller slices, which is what happened at Uber the moment diff count landed on a dashboard. There are narrower caveats, and I won’t pretend I have chased all of them. The comparison is against engineers who already had AI in their IDE, so what it measures is the increment from adding an agent rather than the effect of AI from zero. Four months is not long enough for maintenance cost to turn up. And engineers chose for themselves whether to adopt. None of that makes the study wrong. It is a good measurement of pull request volume, and pull request volume behaves the way lines of code always did. It is easy to report and showcase, if somebody decides that moving it matters. Blind to what the system underneath is doing. What This Actually Means If You Are Leading a Team If you’re six months into your AI investment and your main evidence of ROI is that the output counter went up, whether that counter says lines or pull requests or tasks completed, that should worry you more than it reassures you. You might be measuring the accumulation of future cost and calling it present-day value. Here is what I would look at instead. Cycle time as value reaching production sooner, not just code getting written faster? Those two come apart more often than anyone expects. Rework rate counts if you are fixing more bugs in AI-assisted code? If generated code carries a higher defect rate, your productivity gain is a mirage. Cognitive complexity trend is when your application is getting harder to understand. Tools like SonarQube measure this. If complexity is climbing faster than it did before AI adoption, you’ve got a compounding problem. Developer effort distribution is where the time is actually going. If writing dropped from 20% to 10% of developer time while code review grew from 15% to 30%, you have moved the burden rather than reduced it. None of that is exotic, and it is not just my read. DORA’s 2026 work on the ROI of AI-assisted development arrives at something similar from a different direction. The return on these tools tracks the strength of the engineering system around them rather than the tools themselves. The things that decide it are unglamorous. Whether code review has any slack left in it. Whether people trust the test suite enough to act on a red build, which is the one I have never seen anybody audit. How much sits between a merge and production. Where those are weak, faster generation fills a queue and waits there, and the measurement science is still catching up to the tooling. The Uncomfortable Question Here is what I would ask any engineering leader who reports AI productivity gains based on code output: “If your best engineer spent last week deleting 3,000 lines of AI-generated code and replacing them with 300 lines that do the same thing better, would your dashboard show that as a win or a loss?” On a line count, that week is a catastrophe. On a PR count, it is one merged pull request, which is what a typo fix is worth too. Neither number has any way of seeing what actually happened. If your dashboard shows a loss, you are measuring the wrong thing, and you are rewarding your team for building a larger, slower, more fragile system in exchange for a chart that goes up and to the right. The goal was never more code. It was better systems that deliver business value and are maintained cheaply. Next in this series: Optimizing Benchmark Tasks Instead of Real Delivery Work — why the fact that coding is not the bottleneck makes most AI productivity claims irrelevant to actual delivery speed.

By Gaurav Gaur DZone Core CORE
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do

For the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening. The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified. This is a significant architectural shift. The emerging pattern looks less like: Plain Text prompt → LLM → source code → trust it (the infamous vibe ding pattern) and increasingly like: Plain Text intent ↓ LLM ↓ restricted set of operations ↓ deterministic tools and validation ↓ result Several recent developments in mainstream developer tooling point in exactly this direction. Permission Is Becoming Separate From Intelligence On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations. Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals. That distinction is more profound than it may initially appear. The question is no longer merely: Can the agent perform this operation? It is: Is this agent authorized to perform this operation in this environment? Capability and authority are different things. A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so. A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities. Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering. The Model Is Becoming Replaceable Another development makes the separation even clearer. GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation. This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go. Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear. Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is: Plain Text stable environment stable tools stable constraints stable validation ↑ interchangeable models stable environment stable tools stable constraints stable validation ↑ interchangeable models The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system. Agents Are Separating From Editors VS Code's Agent Host architecture points in another related direction. The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts. Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code: Plain Text human ↓ editor ↓ language server ↓ compiler Agentic development introduces another participant: Plain Text human intent ↓ agent ↓ semantic tools ↓ compiler / tests / environment The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important. A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things. Verification Is Moving From Opinion to Execution A fourth development may ultimately be the most important. GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct. This should sound obvious. We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense. A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system. A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes: Plain Text generate ↓ compile ↓ test ↓ inspect diagnostics ↓ repair ↓ repeat That is substantially more robust than asking a model to inspect its own output and say whether it looks right. The role of the LLM is reasoning. The role of deterministic software remains enforcement. The Interesting Convergence These developments come from different parts of the development stack, but they point toward the same decomposition. An AI programming environment increasingly contains at least four distinct elements: A model that reasons and generatesA vocabulary of operations available to itA capability policy defining which operations it may useDeterministic mechanisms that decide whether the result is valid None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design. What If We Put the Boundary Into the Language? Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted. There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes? This is the idea I have been exploring with an open-source project called BUBAS. BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment. There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary. An order-processing application could, for example, expose operations such as: Plain Text LOAD_ORDER ORDER_TOTAL CUSTOMER_RISK APPROVE REJECT REQUEST_APPROVAL An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed. If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that. Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem. This Is Not a Sandbox Make the distinction carefully. A restricted language does not magically make its host application safe. If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong. The useful guarantee is narrower: Generated business logic can only name operations that the application deliberately made part of its vocabulary. That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment. These operate at different semantic levels. Domain Capabilities Are More Interesting Than Operating-System Capabilities Operating-system permissions are necessary, but business applications eventually need a richer vocabulary. Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful. It still does not answer questions such as: May this program approve an order?May it request approval but not approve directly?Can it read customer risk information?Can it initiate a payment?Can it calculate a premium but not change the underlying policy? Those are domain capabilities. General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction. A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations. We May Be Seeing the New Shape of the AI Programming Stack None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software. But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change. The industry is gradually replacing: Trust the model. with: Give the model precisely defined capabilities and verify what it produces. That is a much more promising engineering principle. BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.

By Peter Verhas DZone Core CORE
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.

By Akhil Madineni DZone Core CORE
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities

An AI agent rarely follows a long plan exactly as first proposed. Tool results expose missing facts, external systems change, policies arrive through human input, and a model may discover that an earlier assumption was wrong. ReAct-style agents were explicitly designed around this interleaving of reasoning, action, observation, and plan updates rather than one immutable plan. The engineering difficulty begins when those actions are durable side effects. In Temporal, a completed Activity is not an intention that can be edited out of a revised plan; its completion and result are part of the Workflow Event History. Replanning therefore has to reconcile a new decision with an already-real past. A Plan Is Intent; History Is Fact The cleanest mental model separates plan state from committed execution state. The plan is provisional: it represents what the agent currently intends to do. Committed state represents facts established by completed Activities, accepted external messages, and other durable events. Temporal stores Workflow history as an append-only sequence and uses that history to reconstruct Workflow state during replay. An ActivityTaskCompleted event contains the serialized Activity result, so a later model decision cannot make that completion disappear. That distinction changes the shape of the agent loop. Instead of asking an LLM to generate a complete plan and then blindly executing it, each replanning call should receive the goal, the latest observations, the prior plan revision, and a normalized set of committed effects. The model may abandon remaining steps, but completed steps should enter the next prompt as constraints on the current world state. Java public AgentResult run(Goal goal) { while (!goalSatisfied(goal, committed)) { Plan plan = planner.replan( new PlanningContext(goal, revision, committed)); for (Action action : plan.readyActions()) { if (alreadySatisfied(action, committed)) { continue; } ActionReceipt receipt = tools.execute(action, action.operationId()); committed.add(receipt); if (receipt.requiresReplan()) { break; } } revision++; } return summarize(committed); } In this shape, planner.replan(...) and tools.execute(...) are Activities, not arbitrary network calls from Workflow code. That matters because Workflow code must remain replay-safe, while non-deterministic operations such as LLM invocations belong behind durable boundaries. Temporal records Activity results in history and returns those recorded results during replay instead of re-running the external call as part of Workflow replay. Temporal’s Java AI integration applies the same principle by executing model calls and external tools through Activities, while deterministic operations may remain inside the Workflow. Replanning Should Consume Effects, Not Merely Step Status A boolean such as completed=true is usually too weak for reconciliation. A completed tool call should return an execution receipt describing the externally relevant postcondition: resource identifiers, amounts, versions, timestamps when materially relevant, and whether a compensating operation exists. The agent then reasons from effects rather than from labels such as “step three succeeded.” Consider a reservation agent that initially plans to reserve inventory, create a shipment, and notify a customer. Inventory reservation succeeds, but shipment creation reports a destination restriction. A revised plan must not schedule another reservation merely because the original plan was discarded. The durable input to replanning should state that a specific reservation already exists and identify it. The next plan might change the shipping method, release the reservation, or escalate for approval, but it should not pretend that the reservation never happened. A second guard is a planning epoch. External input can arrive while a planning Activity is outstanding; Temporal schedules Workflow Tasks when Signals or Updates arrive, and those messages can mutate durable Workflow state. A plan generated from revision 12 should therefore not be committed blindly if the state has advanced to revision 13 before the planner returns. The planner result can carry the revision used as its input, and the Workflow can discard stale output and request a fresh plan. This is optimistic concurrency control applied to agent reasoning rather than database rows. Retries make stable identity equally important. Temporal recommends idempotent Activities because Activities can be retried, and its documentation specifically notes that idempotency keys are appropriate for critical side effects. A plan revision must therefore not double as an operation identity. A reservation that remains semantically the same across replans should retain the same business operation key; a genuinely new reservation should receive a new key. Java public ReservationReceipt reserve(ReservationCommand command) { return inventory.reserve( command.sku(), command.quantity(), command.operationId()); } Here, operationId is passed through to a downstream API that enforces idempotency. The receipt returned to the Workflow becomes durable evidence of the reservation. This also prevents a subtle failure mode in which a model repeats a tool call after losing conversational context even though the Workflow already possesses proof of the earlier effect. Cancellation and Compensation Are Forward Actions A changed plan often creates pressure to “cancel the old step,” but cancellation has a precise boundary. Cancellation can stop work that is still in flight; it cannot retroactively cancel an Activity that has already completed. For long-running Activities, Temporal delivers cancellation cooperatively through Activity heartbeats, so heartbeat configuration determines how promptly an Activity can observe a cancellation request. Once an effect has committed, reconciliation becomes a domain operation. If the effect is reversible, compensation is normally the correct mechanism. Temporal documents the Saga pattern as a sequence of local operations paired with compensating actions, typically executed in reverse order when later work makes earlier effects undesirable. The important semantic point is that compensation creates new history. It does not rewrite old history. A released reservation follows a created reservation; a refund follows a charge; a revocation follows a grant. Java private void reconcile(Plan revisedPlan) { for (ActionReceipt receipt : compensationsInReverseOrder(revisedPlan, committed)) { ActionReceipt reversal = tools.compensate(receipt); committed.add(reversal); } } Not every action has a true inverse. An email cannot be unsent, an external party may already have observed a published event, and a physical operation may have crossed an irreversible boundary. Such effects should be modeled as facts that constrain future planning, not as failures of rollback. The revised plan can issue a correction, create a follow-up notification, or route the case to a human decision, but the historical effect remains part of the state presented to the agent. Durable Agents Need Forward-Only Semantics Replanning also needs to stay distinct from Workflow code versioning. A model changing its runtime plan is ordinary application behavior: a new planning Activity produces a new durable decision after new evidence arrives. Changing deployed Workflow code is different because replay must still produce commands compatible with existing Event History; Temporal provides versioning and patching mechanisms for those code changes. Mixing the two concepts leads to brittle systems in which model variability is handled as deployment variability or, worse, non-deterministic Workflow logic. External corrections fit the same forward-only model. Temporal Signals and Updates can change running Workflow state, and accepted messages become durable inputs that can trigger another planning turn. For very long-running agents, Continue-As-New creates a fresh Event History while carrying forward explicit application state, which makes the committed-effect ledger an important part of the continuation payload rather than transient model memory. The central design rule is simple: an agent may revise intentions at any time, but durable execution never revises facts. Completed Activities should be represented as committed effects with stable identities and useful receipts; in-flight work may be canceled cooperatively; reversible effects may be compensated; irreversible effects must constrain the next decision. Temporal’s history then becomes more than a recovery mechanism. It becomes the authoritative boundary between what the agent merely planned and what the surrounding world has already observed. That boundary allows adaptive AI behavior without sacrificing replay safety, idempotency, auditability, or operational correctness.

By Akhil Madineni DZone Core CORE
Why Incident Response Needs Memory, Not Just Intelligence
Why Incident Response Needs Memory, Not Just Intelligence

Every production incident starts with a simple question: "Has this happened before?" I've lost count of how many incident bridges I've joined where that question came up within the first few minutes. Before anyone proposes restarting a service or rolling back a deployment, someone inevitably starts searching. They look through Slack conversations from previous outages, browse old postmortems, compare dashboards with similar incidents, or dig through runbooks to see whether another team has already solved the same problem. Do you notice what's happening here? The engineers aren't trying to demonstrate how much they know about distributed systems. They're trying to remember. That observation has become increasingly important as AI assistants find their way into engineering organizations. Today's large language models are remarkably good at explaining Kubernetes concepts, debugging stack traces, writing SQL queries, or summarizing log files. Those capabilities are valuable, but they only solve part of the problem. During an incident, reasoning is rarely the bottleneck. Finding the right context is. The engineer who resolves an outage the fastest isn't always the one with the deepest theoretical knowledge. More often, it's the engineer who remembers that a similar issue occurred eight months ago after a database failover, or who knows that a particular service has historically exhibited the same failure pattern after specific deployment changes. Experience is a form of memory. If we want AI to become a trusted operational partner instead of just another chatbot, we need to think about memory as carefully as we think about intelligence. Intelligence Answers Questions. Memory Solves Problems. Large language models are exceptional at answering questions because they've been trained on an enormous amount of public knowledge. Ask an LLM to explain consensus algorithms, Kubernetes scheduling, or distributed tracing, and you'll probably receive a detailed, technically accurate explanation within seconds. Production incidents, however, ask very different questions. Instead of asking: What is Kubernetes? Engineers ask: Why did our Kubernetes cluster start failing after yesterday's deployment?Has this service failed in the same way before?Which runbook actually worked the last time?Who owns this dependency today?What changed in the last hour that could explain this behavior? Those aren't questions about computer science. They're questions about organizational memory. The answers don't exist in a foundation model's training data because they're unique to every organization. They live in deployment histories, incident timelines, internal documentation, architecture decisions, monitoring dashboards, chat conversations, and postmortems accumulated over years of operating software. An AI that understands distributed systems but lacks access to this operational history is like an experienced consultant joining your incident bridge for the first time. It may offer useful suggestions, but it doesn't know your environment, your systems, or your team's accumulated experience. That's why intelligence alone isn't enough. Every Incident Is a Search Problem One pattern I've noticed is that incident response often looks less like debugging and more like information retrieval. Think about what engineers actually do during the first fifteen minutes of a major incident. One person opens dashboards to identify where the failure started. Another compares the current deployment with the previous version. Someone searches Slack for keywords that resemble the current symptoms. Another engineer opens the last postmortem for the affected service. Meanwhile, the incident commander tries to understand which teams need to be involved. None of these activities involve writing complex algorithms. They're all attempts to reconstruct context. If you mapped the engineer's workflow, it might look something like this: Alert → Metrics → Logs → Deployment History → Previous Incidents → Runbooks → Slack Discussions → Architecture Documentation → Decision The common thread is that engineers are constantly retrieving information before making decisions. That retrieval process is exactly where AI can provide the most value—not by replacing engineering judgment, but by dramatically reducing the time required to gather relevant context. Not All Memory Is the Same When we talk about memory in AI, it's easy to think only about conversation history or a vector database. In practice, incident response depends on several different kinds of memory, each answering a different set of questions. Incident memory This is the collective history of operational failures. Previous incidents, timelines, root causes, postmortems, and lessons learned all fall into this category. During an outage, one of the first questions engineers ask is whether they've seen the problem before. An AI that can retrieve similar incidents and explain how they were resolved immediately provides value because it shortens the investigation. Operational memory Runbooks, playbooks, escalation procedures, and service ownership represent another form of memory. These artifacts capture how an organization expects engineers to respond under different circumstances. Instead of generating a generic remediation plan, an AI can recommend the procedure that has already been validated by the organization. Infrastructure memory Production systems change constantly. Deployments, feature flags, infrastructure updates, configuration changes, and dependency upgrades all influence system behavior. Understanding what changed recently is often more useful than understanding how a technology works in theory. Organizational memory Some of the most valuable operational knowledge never reaches formal documentation. Engineers discuss recurring issues in Slack, record architectural decisions in design documents, and exchange troubleshooting tips during retrospectives. Over time, this becomes institutional knowledge that experienced engineers rely on instinctively. AI should be able to surface that knowledge instead of forcing every engineer to rediscover it. Memory Changes the Quality of Recommendations Imagine two AI assistants responding to the same latency alert. The first assistant says: CPU utilization is high. Consider restarting the service. It's not necessarily wrong, but it's also not particularly helpful. Now imagine a second assistant with access to organizational memory says: A similar incident occurred three months ago after deployment version 6.4. During that incident, restarting the service temporarily reduced latency, but the underlying cause was an inefficient database query introduced by the deployment. The query was reverted, and latency returned to normal within six minutes. A deployment with similar changes occurred eighteen minutes before the current alert. I recommend validating query performance before restarting the service. Neither assistant is more intelligent in the traditional sense. The second assistant is simply making better use of memory. That additional context changes the recommendation from a generic suggestion into operational guidance grounded in the organization's own experience. Building Memory Into AI Systems Memory isn't a single database or a single technology. It's an architectural capability that combines multiple sources of operational knowledge into a coherent context for reasoning. A production-ready incident assistant might continuously ingest information from observability platforms, deployment pipelines, service catalogs, incident management systems, internal documentation, and communication channels. Rather than asking engineers to manually gather information from each source, the AI assembles the relevant context before generating a recommendation. The language model is still responsible for reasoning, summarization, and communication. The memory layer ensures that reasoning is grounded in facts that are specific to the organization rather than generic patterns learned during training. In many ways, this mirrors how experienced engineers work. They don't solve incidents by relying only on theoretical knowledge. They combine technical understanding with years of accumulated operational experience. That's exactly the capability our AI systems should emulate. Final Thoughts As large language models continue to improve, it's tempting to believe that more intelligence alone will solve the challenges of operational AI. My experience suggests otherwise. The most effective incident response systems aren't necessarily the ones with the largest models or the most sophisticated prompts. They're the ones that help engineers remember. They surface the right runbook, identify the last time a service failed in the same way, highlight the deployment that introduced the problem, and connect today's symptoms with yesterday's lessons. In other words, they make organizational experience accessible when it's needed most. Incident response has always been a combination of reasoning and memory. AI has made remarkable progress on the first half of that equation. The next step isn't simply building smarter models—it's building systems that remember.

By Akshay Pratinav
Software Quality Habits and AI
Software Quality Habits and AI

Software quality is about habits. The quality of our releases depends on the quality of our habits. Nobody ships a defect-riddled release because they didn't want quality. They ship it probably because the small daily behaviors that would have prevented it — the small habits — quietly stopped happening. One skipped test, one PR approved without being read, one “we'll fix it later” at a time. This article makes a quick introduction to those habits, small and large. It explains how small habits become large over time and how good and bad habits compound. I also discuss a specific danger I think we're underestimating: that AI, used without understanding and judgment, doesn't just fail to build good-quality habits. It can actively erode the ones that we may already have. Small Habits Small (or atomic) habits are small enough to survive a bad day and cheap enough to repeat without thinking about it. None of them is impressive in isolation. That's exactly why they're the unit that everything else is built from. Writing and testing: Writing the failing test before writing the fix, not afterChecking the null case, the empty list, the empty string, the zero, the timeoutWriting one assertion that actually checks behaviorDeleting a test that no longer tests anything meaningful instead of leaving it as decorationRunning the full suite locally before pushing, not just the file you touchedAdding a regression test the same day a bug is fixed, not “when there's time” Code review and collaboration: Actually reading a diff line by line before approving itLeaving a comment when you don't understand something, instead of rubber-stamping itAsking “what happens if this call fails” on every PR that adds a network or disk callRequesting a second reviewer on anything touching auth, payments, or data migration, without being told toReviewing your own diff once before asking anyone else to Naming, structure, and documentation: Naming things so the next person doesn't need you to explain themUpdating the doc or README in the same PR as the code change, not in a follow-up that never comesWriting a one-line commit message that says why, not just whatDeleting dead code the moment you notice it, instead of leaving it “in case”Logging enough context at the point of failure that you won't need to reproduce it just to read a log Everyday discipline: Reading the error message fully before searching for itReproducing a bug before claiming to have fixed itFlagging a flaky test the first time you see it, instead of re-running until it passesSaying “I don't know why this works” out loud instead of merging it anyway Large Habits Large habits need sustained effort and organizational will, not just individual discipline. They're expensive to install and easy to let quietly decay. I haven't worked with many teams that actually have the large habits that they think they have. A test suite that's trusted enough that a red build actually stops a merge, every time, with no manual override cultureA blameless postmortem process that produces real changes, not just a document nobody rereadsA quality gate in the release pipeline that blocks a release rather than getting routed around under deadline pressureAn onboarding program that transmits testing and review culture to new hires, not just repo access and a Slack inviteArchitecture reviews that ask “how will this fail” and “what happens at 10x load” before “does this work”A living catalog of architectural decisions and the reasoning behind them, kept current enough that people actually consult itA production monitoring and alerting setup tuned enough that on-call trusts the alerts instead of muting themA deprecation and technical-debt process with real budget and real deadlinesChaos engineering or systematic failure-injection practiced routinely, not once after a bad incidentA culture where raising a quality concern is rewarded even when it slows a release down How Small Habits Become Large A large habit is what a small habit looks like after a few hundred repetitions and a bit of organizational scaffolding around it. A few concrete examples: Writing one unit test with every change, consistently, for a year, is what a real unit test suite is made of. Unit test suites often accumulate, PR by PR, from the atomic habit of not merging untested logic. A test suite guards against regressions once many individual tests are maintained and are trusted enough so that a failure blocks a merge.Leaving one honest review comment when something is unclear, repeated across hundreds of PRs, is what eventually produces the right review culture. A culture where junior engineers feel safe asking questions and senior engineers feel obligated to answer them carefully. No policy document creates that; it's the residue of a habit practiced long enough to become the norm.Writing down why a bug happened, every time, in a shared place, is what a real blameless postmortem process is made of. The first few write-ups are just notes. After enough of them accumulate, cross-referenced and occasionally reread, they become an institutional memory the team consults before making the same category of mistake twice.Updating a doc or writing a short note about a decision each time a decision is made is what eventually becomes a living architectural record. This is something that new hires can read to understand not just what the system does, but why it's shaped the way it is.Flagging one flaky test instead of re-running it is what, multiplied across a team and a year, is the difference between a CI pipeline people trust and one people route around. The large habit — a quality gate that actually informs us about releases — cannot exist unless enough individuals have consistently practiced the small habit. The direction of causation matters here. You cannot batch-install a large habit. You can write the policy, buy the tool, mandate the process — but if the atomic habits underneath it aren't being practiced by individuals, the large habit is just a costume. The postmortem process without honest small write-ups is meaningless. The quality gate without individuals who trust failing checks is a formality people learn to bypass. Large habits are the compounded, aggregated shape of thousands of small habits; you build them from the bottom up, or you don't really have them at all. Habits Compound — In Both Directions A team that writes one more edge-case test per PR than it used to, for two years straight, ends up somewhere completely different from a team that writes one fewer. Neither team can point to the day quality became a strength or a liability. It's a rounding error every day and a number of dissatisfied customers over a year. The bad-habit version compounds just as reliably. “We'll add the test later” becomes later never comes. That's how test suites that nobody fully trusts are built. That's also how red CI that gets ignored becomes a quality gate that's causing problems instead of solving them. The organization now believes it has quality controls it doesn't actually have — and that belief gap is dangerous. Where AI Genuinely Helps Currently, AI is legitimately good at collapsing friction. Writing code and drafting test scaffolding or suggesting the edge case a tired engineer misses at 6 p.m. Summarizing a diff so a reviewer actually reads it instead of skimming, flagging the null case nobody handled — these are atomic-habit amplifiers. They make good habits cheaper to perform, and cheap good habits are often the ones that survive. As AI keeps improving, I expect it to be good at more and more aspects of the SDLC. Where It Becomes Dangerous “AI is good and bad” is a cliché I don't want to hide behind, so here's the specific mechanism. The danger isn't that AI writes bad code or hallucinates an API. It's that when AI fully closes a loop — writes the fix, writes a test that passes, resolves the incident — it can quietly remove the exact process that used to build the judgment we rely on to know when something is fragile. Before AI: an engineer hits a bug, doesn't understand it immediately, tries something, fails, forms a theory, tries again, and eventually understands the system well enough to fix it. More importantly, well enough to recognize the next three places the same class of bug will show up. The fix is important, but it is not the most valuable output of that process. The mental model built while failing is the most valuable output. When AI closes the loop opaquely, even if the bug is really fixed, the mental model may never get built. Engineers will lose knowledge while their judgment and critical thinking weaken. The ability to map incidents to root causes, to understand a system in more depth as we analyze it again and again — all that is gone if we simply accept an AI output without understanding. Terence Tao on AI in Mathematics Fields medalist Terence Tao raised a similar argument about mathematics. His claim isn't that machine-generated proofs are worthless but that the field of mathematics is stagnating because it is entering a crisis of value and understanding. Good, fruitful open problems are mined in a non-renewable fashion, Tao mentions. He goes on to explain: "In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained." In essence, Tao isn’t objecting to AI’s capability; he is objecting to AI’s opacity. Using the Navier–Stokes regularity problem as his example, Tao described a scenario where an AI system runs an entire search process privately and simply hands back a finished proof. The problem may get solved, but the field gains almost nothing. And this is because the value of a hard open problem is not in the answer itself. The value lies in the decades of people's failed attempts: the tools that we had to invent along the way, the adjacent structure that we mapped while stuck, the wrong turns that became new subfields on their own. Solve a problem transparently, and the field grows by understanding and learning. Solve it as a sealed black box, and you collect the answer but skip the part that was actually producing the mathematics. Let's think about that for a moment from an engineering perspective. Aren't open problems, and how well we understand them, one of the main reasons for engineering innovation? The struggle to keep distributed systems consistent and available produced the CAP theorem. By understanding the CAP theorem, consensus protocols like Paxos and Raft now exist in reliable infrastructure. Trying to run software at a scale that no team could reason about manually gave rise to container orchestration and service meshes. Such open problems gave rise to site reliability engineering as a discipline in its own right. The difficulty of trusting code that could never be exhaustively tested gave birth to property-based testing, fuzzying, and formal verification tools like TLA+. Realizing that testing everything is often meaningless and inefficient (if possible at all) shaped software testing as a discipline. Chaos engineering exists because a team treated “our systems fail in ways we can not predict” as a problem worth living inside rather than a problem to be solved instantly without understanding the solution. Open Problems And Software Quality Our version of “open problems” is production incidents, hard-to-reproduce bugs, and the uncomfortable stretches where nobody's quite sure why the system behaves the way it does. Historically, those gaps close the slow way: someone traces the failure, builds a mental model, writes it up, and the team's collective judgment gets a little better. That's our version of the “adjacent structure mapped while stuck” — debugging notes, a workaround that became a design pattern, the engineer who six months later says “oh, I've seen this shape before.” An AI that resolves an incident by generating a fix nobody on the team fully understands solves the problem in the narrowest possible sense and gives back nothing else. Multiply that across a year of incidents, and you get an organization that ships, that's green on every dashboard, and that has steadily lost the tribal knowledge that used to make it resilient. The loss won't show up as a regression in this quarter's metrics. It shows up the next time a hard problem appears, and nobody in the room has the instinct for it — because the instinct was never built. It was skipped. To keep the loop intact, we could: Treat an AI-generated fix like a junior engineer's PR, not a senior engineer's judgment call — someone accountable still explains, in their own words, why it works before it mergesDon't let AI's speed quietly cancel the postmortem — if it was a real incident, someone still writes down what happened, even if the fix took ten minutes instead of two daysUse AI to compress the mechanical part of the habit loop — scaffolding, boilerplate — but keep the diagnostic “why did this actually fail” a human exercise, at least until an AI's explanation has been checked and not just found plausibleWatch for the belief gap: a green pipeline built partly on unexamined AI fixes can look identical to one built on understood fixes, until it doesn'tWhen onboarding, keep exposing new engineers to the failed attempts, not only the final AI-polished diff — the failed attempts are still where judgment gets transmitted Wrapping Up Identity is built by repeated evidence because in software we are what we repeat. We are our habits. We become someone who doesn't skip tests by repeatedly not skipping tests. Quality cultures work the same way at the team level. One repetition at a time, small habits accumulating into large ones. AI can make each of those repetitions cheaper, and that's genuinely valuable. What it must not be allowed to do is to remove our learnings from repetitions themselves and hand back only the outcome. As mathematics loses something specific and non-recoverable when a hard problem is solved opaquely, software quality is exposed to the same risk — on a much shorter feedback cycle, in production, right now.

By Stelios Manioudakis DZone Core CORE
Can Your Team Name the Work It Already Runs With AI?
Can Your Team Name the Work It Already Runs With AI?

TL;DR: The AI Workflow Inventory Finalizes the A3 Delegation System You probably know your own AI shortcuts: the colleague who drafts the stakeholder update, the nightly job somebody set up before they left, the interview notes that go through a model on demand. However, I want to challenge you: that perceived knowledge can create false confidence, because it feels like knowing the team’s way of working with AI, when it is still only a diffuse understanding of the practice. Ask the team to combine those individual accounts into one list, and you may discover how little of the whole anyone can see. The AI Workflow Inventory is the artifact for that list: one row per recurring task, refined into provisional task classes to prepare the team’s next AI delegation decisions. Disclaimer: I read Charniak/McDermott’s book on “Artificial Intelligence” decades ago; of course, I use AI for research, translation, proofreading, challenging story arcs and article structures, and summarization. It is a production tool, not a substitute for thinking. Your Own Shortcuts Are the Easy Part When I published the AI Delegation Lifecycle in June 2026, the point before stage 1 was deliberately left without an artifact: the team was supposed to know which work it does with a model, at what frequency, and at what stakes, which I call a forensic analysis of your own workflow. However, as I quickly discovered, this assumption that a team could produce the analysis from memory was wrong, and updating the student reference for the A3 Delegation System earlier this month showed me its design flaw: The Routing Policy assigns an AI model tier per task class,The AI Definition of Done is written per task class (never per task), andThe Delegation Audit opens with a walkthrough of what the team has handed over. All three artifacts consume task classes, but none of the six stages of the A3 Delegation System produces them. The AI Working Agreement is the exception on the other side of the lifecycle: it records the team’s defaults and boundaries and can exist before the inventory is complete. The gap is easy to miss from the inside, because every individual on the team can answer for their own use, and nobody can answer for the sum: who else drafts with a model, which outputs leave the team that way, and which jobs run unattended. That sum, useful, undocumented, and unowned, is what I call “AI Debt,” and it works right up to the moment someone asks who is responsible for it. What One Row Looks Like in the AI Workflow Inventory The AI Workflow Inventory is a one-page canvas with an identity strip (team, owner, date, review date), a panel on where to look, a panel with five refinement rules, and the inventory table itself. Let me show you an example from a Scrum team: Number: I-01Task: Transcribe photos of Retrospective sticky notesProcess: Retrospective facilitationTask class: Transcriptions for internal useHow often: Per SprintWho runs it: Scrum MasterTool: GPT-5.6 SolOutput goes to: Stays in the teamData that enters the model: Photos of handwritten notes and names. The task is written as verb and object, the class is named by output and audience, and “output goes to” has three values: stays in the team, leaves the team, or customer-facing. The number travels to the Workflow Card later, so the row and the respective decision can be matched months apart. Two design choices carry most of the weight: The person and the tool sit in separate columns because people change and tools change while the task stays; a field that says “Anna, ChatGPT” makes it harder to update either one when Anna moves teams or the license changes.The last two fields ask different questions that teams routinely conflate: where the output goes tells you about the stakes on the way out, whereas what enters the model tells you about the exposure on the way in, and an output that “stays in the team” can still have been produced from photographs with names on them. The entry in the data field is a fact to review; it is not permission to keep uploading the names. What the canvas does not have is a life cycle stage or status column, and the omission is deliberate. The inventory records what happens today; the poster and the Workflow Cards show where a workflow stands in the A3 Delegation Life Cycle, and a retired workflow is a row in the Re-classification Log, identified by the same inventory number. A stage column would turn the pre-decision list into a second, competing status board, and agile practitioners already know how that ends from the Product Backlog in Jira and the “real” one in a spreadsheet debate. The AI Workflow Inventory takes work to maintain, but it reduces repeated documentation and decision-making: The team captures each recurring task once.Compatible tasks share one review standard and one routing decision instead of getting one each.The inventory number connects a task to every later decision about it.The life cycle status stays on the poster and in the records that already exist. Update the inventory when the work changes or when the team discovers, corrects, or refines what it has recorded. (I know, it sounds like a database schema, and I started preliminary work on how to design an application for the A3 Delegation System.) Where to Look The canvas names three sources for AI workflows, and the 30-day window matters because memory is a poor inventory tool, while last month’s outputs are evidence: Your meetings come first: which outputs of the events that repeat (planning, pipeline or campaign reviews, stand-ups, monthly reporting, Retrospectives, etc.) did a model draft, summarize, transcribe, or translate?Then your tools: each colleague looks at their own chat history from the last 30 days and brings the prompts, applied Skills, or agent tasks that come back every week or every Sprint, who wrote them, and how often they were rewritten; this is a self-report, and it has to be, since nobody should be reading anyone else’s conversations.Then your outbox: the emails, reports, tickets, and documents that left the team, and the question of which of them started as a model draft. Thirty days is the starting window. Once the obvious rows are on the page, ask about the work no chat history shows: the jobs that run unattended, the quarterly report, the thing that only happens at release time. A nightly test run, such as the one in the example below, may leave no trace in anyone’s recent chat history. The line that decides the session is the last one in that panel: list the shortcuts nobody approved as well, because that is where the AI debt sits. If people read the inventory as a check on whether they followed the rules, those shortcuts stay invisible, and the team ends up with a list that looks tidy and describes a team that does not exist. Therefore, going first with a shortcut of your own helps, as does explaining, before the first row, that recording a task does not approve its use: the session starts with understanding what happens today, and concerns that need immediate attention under the team’s existing boundaries will still be addressed. Provisional Task Classes and the Grouping Test The grouping rule creates a puzzle that a careful reader will spot: it asks whether two tasks share the same AI Definition of Done, the same reviewer, and the same model tier, and those are decisions for later stages of the AI Delegation lifecycle, which the team has not reached. The resolution is that classes in the first session are provisional. Name them by output and audience, test the grouping against whatever review standards and routing decisions the team already has, and where those are missing, leave the grouping provisional and revisit it as the team works through the stages. An AI workflow inventory can carry uncertainty; what it must not carry is a settled-looking decision nobody made. The five rules on the AI Workflow Inventory canvas are short: Capture first, judge later: The A3 decision comes in stage 1, after the row exists, and a row approves nothing.One row per recurring task: Four prompt revisions used for the same weekly update are one task, and the tool goes in its own column.Name the class by output and audience: “Transcriptions for internal use,” “status communication leaving the company.” A class name that fits any task (“AI support”) is too wide to tell anyone which standard applies. (That naming approach is not different from coding.)Same standard, same reviewer, same tier, one class: When one of the three differs, split the class.Every class goes to stage 1; every audit refines the list: A class without an A3 decision is AI debt with a date on it. Three failure patterns are worth watching for: the sanctioned-use list, where only approved use gets written down, and your existing AI debt stays invisible; the prompt or Skill catalog, with a row per prompt/Skill instead of per task, so the classes never form; and the one-time census, filled once and never refined, so that six months later the audit walks a stale list and everyone nods at rows that no longer run. Example: Four Rows From a Scrum Team Here is a full first pass from a fictitious Scrum team of seven: Transcribe photos of Retrospective sticky notes (Scrum Master, GPT-5.6 Sol, stays in the team; input: handwritten notes with names): Internal output can still involve sensitive input.Rewrite customer feedback tickets into Product Backlog item drafts (class: requirement drafts for internal use): A concrete task inside a reusable class.Draft acceptance criteria for new items (same class): A second task that may share that class, as long as reviewer, standard, and tier stay the same.Rewrite interview notes into the hiring feedback template (on demand, Scrum Master, leaves the team; input: candidate names): Informal use that an approved-use-only list would have missed. The last row is on the list only because capture came first; under a rule that lists only approved use, the person running it would have left it out, and the team would have gone into A3’s stage 1 without knowing candidate names were entering a model on demand. From Row to Workflow Card Each decided row of the AI Workflow Inventory becomes a Workflow Card within the A3 Delegation System: workflow, task class, the A3 category, and later the tier, the owner, and the audit dates; the card goes on the poster where the workflow stands today. The word “decided” matters: the row exists before the A3 “Assist-Automate-Avoid” decision; the card exists after it. A row without a card means one of two things, and the remedy differs: either nobody has decided, in which case the class goes to the A3 Framework next, or somebody decided and never recorded it, in which case the team confirms the decision and writes it down. Ask which it is before doing either. For the transcription row, that means: the AI Workflow Inventory gives it a number, a task, and a provisional class; A3’s stage 1 decides whether it is Assist, Automate, or Avoid, and knowing that GPT-5.6 Sol runs it today settles nothing; stage 3 settles the owner, and the Scrum Master who runs it this Sprint does not become the owner by default. Both decisions go on the card. Your First 60 Minutes with the AI Workflow Inventory Book the session, aim for eight rows, and treat it as a first pass. One way to split the time: State the purpose and the boundary (capture, not approval): 10 minutesWalk the three sources, including unapproved shortcuts and unattended jobs: 20 minutesForm provisional task classes with the grouping test: 20 minutesName the inventory owner and the review date, and agree on the next step: 10 minutes The next step is A3’s stage 1: take every class to the A3 decision with the people who run the tasks in the room. If a row needs attention under the team’s existing boundaries before that (candidate names in a model, say), address it; “capture first” gives the decision its own step and does not postpone it. At each Delegation Audit, the list gets two kinds of maintenance. Add the rows that appeared since the last audit, and reconcile the existing ones: does the task still run, does the same person run it, with the same tool, on the same inputs, for the same audience? When you retire a task, record that decision in the Re-classification Log under the same inventory number, and keep the link to its original inventory record. Conclusion Which recurring use would your team’s current list miss? If your team cannot describe where AI already contributes to recurring work, start with a 60-minute AI Workflow Inventory session. The list gives you a shared basis for A3 decisions and helps you find the workflows that individual recollection might miss; then choose where in the A3 Delegation System to apply the method first.

By Stefan Wolpers DZone Core CORE
Locking Down the Enterprise: Data Security Patterns for AI Integrations
Locking Down the Enterprise: Data Security Patterns for AI Integrations

Every enterprise conversation about artificial intelligence eventually arrives at the same uncomfortable question: what happens to our data once it leaves our perimeter? Whether you are wiring a large language model into a claims processing pipeline, standing up a retrieval-augmented generation (RAG) system for internal knowledge search, or letting an agentic workflow take autonomous actions against production systems, the answer determines whether your AI initiative becomes a competitive advantage or a compliance incident waiting to happen. This article lays out a layered, defense-in-depth approach to securing enterprise data across the AI lifecycle — from data classification and access control, through transit and storage, vendor contracts, prompt hygiene, architecture patterns, and compliance mapping. It is written for architects, technical leads, and engineering managers who are past the "should we use AI" conversation and are now living in the "how do we do this safely at scale" reality. The guidance here is deliberately vendor-agnostic and framework-agnostic. The specific tools you choose — which cloud, which model provider, which vector database — will vary. The principles will not. 1. Why AI Changes the Data Security Calculus Traditional application security assumes a relatively closed loop: your code, your database, your network boundary. Data moves through defined pathways, and you can reason about every hop. AI systems break several of these assumptions at once. The boundary is porous by design. A large language model call is, functionally, an API request to a third party — even when that third party is a trusted enterprise vendor. Every prompt is an egress point. Every completion is an ingress point. Unlike a traditional API integration where the schema is fixed and the payload is structured, prompts are free text, which makes it far easier for sensitive data to slip in unnoticed. The system can be instructed by its input. In a conventional application, data and instructions are cleanly separated — SQL injection exists precisely because that separation sometimes breaks down, and we have spent two decades building defenses against it. In an AI system, the model's instructions and the data it processes often occupy the same channel: natural language. This is the root cause of prompt injection, and it means the data itself can become an attack vector, not just a target. The system can act, not just answer. Agentic AI — systems that call tools, write files, send emails, or modify records — collapses the distinction between "the AI leaked data" and "the AI did something harmful with data." A single compromised or manipulated agent can chain read access into write access, and write access into external communication, all within one interaction. The data footprint compounds. Vector embeddings, cached completions, fine-tuning datasets, evaluation logs, and conversation histories all represent new copies of your sensitive data, sitting in new places, governed by new retention rules that your existing DLP and archiving policies were never built to see. None of this means AI is unsafe to deploy in the enterprise. It means the security model has to be designed deliberately, layer by layer, rather than inherited by default from your existing application security posture. 2. Data Governance and Classification Everything downstream depends on getting this layer right first. If you don't know what data you have and how sensitive it is, no amount of encryption or access control will save you from sending the wrong thing to the wrong place. Classify Before You Integrate Before any AI system touches production data, classify it into tiers. A simple, workable scheme: Public – marketing content, published documentation, anything already externally visible.Internal – operational data with no direct regulatory exposure, but not meant for public release.Confidential – customer PII, employee data, financial figures, strategic plans.Restricted – regulated data categories: PHI under HIPAA, PCI cardholder data, biometric data, data covered by state insurance data security laws, or anything under a specific contractual non-disclosure obligation. Each tier should carry an explicit, written policy on whether and how it may be used with AI systems — including which AI systems (internal, VPC-isolated, or public API) and under what redaction or tokenization requirements. Data Minimization Is a Design Constraint, Not an Afterthought The single most effective control available to you is simply not sending data you don't need to send. This sounds obvious and is routinely ignored under deadline pressure. Practical patterns: Field-level scoping: If a prompt needs a claim status and adjuster name, don't serialize the entire claim record into context. Query for and pass only those fields.Row-level scoping in RAG: Retrieval pipelines should filter at the query layer (based on the requesting user's entitlements) before documents ever reach the context window, not rely on the model to "know" what it shouldn't discuss.Aggregate over raw where possible: If the use case is trend analysis, send aggregated statistics rather than the underlying raw records. Understand Retention and Training-Use Terms This is the question every enterprise security review should ask first, and the one most often skipped: does the AI provider retain my inputs and outputs, and are they used to train or improve models? Enterprise API tiers from major providers typically differ meaningfully from consumer-facing chat products on this point — enterprise agreements commonly include zero data retention (ZDR) options and explicit commitments that customer data is not used for model training. Consumer tiers, free tiers, and browser extensions are a different story and should be treated as such in your acceptable use policy. Don't assume; read the actual data processing terms for the specific tier and product you're using, and get it in writing. Data Lineage for AI-Touched Data Once data has passed through an AI system, it has effectively been transformed and potentially recombined. Maintain lineage records: which source systems fed which prompts, which model version processed them, and where the outputs were stored or acted upon. This becomes essential later for both incident response and regulatory audit. 3. Access Control for AI Systems Treat AI Service Accounts Like Any Other Privileged Identity An AI agent or pipeline that calls internal APIs is a service account. It should be provisioned, reviewed, and revoked exactly like any other service account — with the added scrutiny that its "instructions" can be influenced by untrusted input in ways a traditional service account's code path cannot. Least privilege, scoped by task: A customer-support chatbot that looks up order status needs read access to an orders API — not write access, not access to the full customer database, not admin scopes "just in case."Short-lived credentials: Prefer short-lived, automatically rotated tokens (OAuth client-credentials flows, workload identity federation) over long-lived static API keys.Per-tenant isolation: In multi-tenant SaaS or multi-client environments, ensure the AI system cannot cross tenant boundaries even if a prompt attempts to coax it into doing so — enforce this at the data access layer, not the prompt layer. Enforce Human Entitlements Downstream of the Model, Not Just Upstream A common and dangerous mistake: building a RAG or agentic system where the AI service account has broad access "for flexibility," and relying on the system prompt to tell the model which documents the current user is allowed to see. Prompts are not an access control mechanism. If the underlying retrieval or tool-calling layer can technically reach a document or record, a sufficiently motivated (or simply unlucky) input can potentially surface it. The correct pattern is to filter at the data layer using the actual requesting user's entitlements — row-level security in the database, document ACLs in the retrieval index, and scoped API tokens minted per-request based on the authenticated user, not the service account. Role-Based Access for AI Outputs Re-Entering the System When an agentic workflow's output writes back into production — updating a record, sending a notification, filing a claim note — that write should pass through the same RBAC and validation layer a human-initiated write would. Do not grant an AI agent a privileged bypass "because it's automated." Automation is exactly when you want the guardrails to be strongest, because there is no human in the loop to notice something is wrong before it happens. 4. Secrets and Credential Management This deserves its own section because AI systems introduce new and easy-to-miss places for secrets to leak. Never hardcode credentials in prompts, system prompts, or agent configuration files. It is tempting to embed an API key directly in a tool definition during a proof of concept. That habit does not survive contact with production.Never let secrets end up in agent memory or long-term conversation history. If your architecture includes persistent memory for an agent, explicitly exclude credential material, and audit what actually gets written to that memory store.Use a dedicated secrets manager — Azure Key Vault, AWS Secrets Manager, HashiCorp Vault, or your platform equivalent — and have the AI orchestration layer fetch credentials at call time rather than holding them statically.Rotate aggressively for anything touched by an AI pipeline. Given that prompts and tool definitions are more likely to be copy-pasted into documentation, shared in Slack for debugging, or logged verbosely during development, treat any credential that has been anywhere near an AI pipeline as higher-risk and rotate on a shorter cycle.Watch your logs. Verbose request/response logging — common during AI development for debugging prompts — is a frequent, unglamorous source of credential leakage. Redact before logging, not after. 5. Data in Transit and at Rest The fundamentals here are not AI-specific, but they are easy to underinvest in because AI integrations often move fast and get treated as "just another API call." TLS everywhere, including between internal orchestration services and the AI provider, and between internal services and any vector database or cache.Encrypt at rest, including: The primary data stores feeding your RAG pipeline.Vector embeddings themselves. Embeddings are not inherently anonymous — depending on the embedding model and dimensionality, source text can sometimes be partially reconstructed from vectors, so treat an embedding store with the same sensitivity as the source documents.Prompt and completion logs.Any cached responses (semantic caching layers are increasingly common for cost control and latency, and they represent another copy of potentially sensitive data at rest).Encrypt backups of all of the above, and include them explicitly in your data retention and destruction policies — a backup snapshot of a vector database is a backup of your confidential documents. 6. Vendor and Contractual Controls Technical controls only get you so far if the underlying contract with your AI provider doesn't back them up. Zero Data Retention Agreements Where available, negotiate zero data retention (ZDR) terms — an explicit commitment that request payloads are not retained beyond the time needed to serve the response, and are not logged, cached, or used for any secondary purpose. This is increasingly available as a contractual option from major enterprise AI providers and should be a standard line item in procurement for any AI vendor touching confidential or restricted data. Data Processing Agreements A proper Data Processing Agreement (DPA) should cover: Purpose limitation (data used only to provide the contracted service).Sub-processor disclosure (who else touches your data downstream of the primary vendor).Data residency commitments (does data ever leave a specific geographic or regulatory jurisdiction).Breach notification timelines.Audit rights. Enterprise Tier Versus Consumer/Shared Infrastructure Confirm explicitly whether you are on infrastructure that is logically or physically isolated from other customers, versus a shared multi-tenant consumer product. Ask directly: is my data ever used to train models that serve other customers? Is there any possibility of cross-tenant data mixing in caching or logging layers? Get the answer in the contract, not just in a sales conversation. Vendor Security Posture Review Standard vendor risk management practice applies, but with AI-specific questions added to the questionnaire: What is the model provider's own subprocessor chain?How is prompt injection or jailbreak resistance tested and monitored on their side?What certifications do they hold (SOC 2 Type II, ISO 27001, ISO 42001 for AI management systems specifically)?What is their incident response commitment and SLA for a security event affecting your data? 7. Prompt and Output Hygiene Treat Untrusted Content as Untrusted, Even Inside a Prompt Prompt injection is the AI-era equivalent of injection attacks in traditional application security, and it deserves the same rigor. Any content that originates outside your organization's direct control — an email, an uploaded document, a web page fetched by a tool, a third-party API response — should be treated as untrusted input, not as trusted instructions, even when it is concatenated into the same prompt as your system instructions. Practical mitigations: Clear structural separation between system instructions, trusted context, and untrusted content, using explicit delimiters and, where the platform supports it, distinct message roles.Instruction-following boundaries: Explicitly instruct the model that content within untrusted blocks should be treated as data to analyze, not as commands to follow — and validate this behavior in testing with adversarial inputs, not just happy-path examples.Least-privilege tool access during untrusted content processing: If an agent is currently processing an untrusted document, don't give it simultaneous access to high-privilege tools (sending email, executing code, modifying records) without a human confirmation step in between. Output Validation Before Action Any AI output that will be displayed to a user, stored in a system of record, or used to trigger a downstream action should pass through validation: Schema validation for structured outputs (if you asked for JSON, validate it actually conforms before using it).PII/sensitive-data scanning on outputs, not just inputs — a model can sometimes surface data it was never explicitly asked to reveal, particularly in RAG systems with imperfect retrieval filtering.Action confirmation gates for anything irreversible or high-impact — sending external communications, financial transactions, deleting records — even in a fully agentic workflow. A "dry run" or human-approval step for a defined set of high-risk action types is a small latency cost for a large risk reduction. PII Redaction Pipelines For any workload where the AI system doesn't strictly need to see PII to do its job, run a redaction or tokenization pass before the data reaches the prompt, and a re-hydration pass on the output if needed. This is particularly relevant when using third-party or shared-infrastructure LLM endpoints for tasks like summarization or classification, where the specific identity behind the data is often irrelevant to the task itself.

By Balaji Venkatasubramaniyar DZone Core CORE
How Multi-Agent Systems can replace most of Manual ML Validation decisions - The Karpathy Loop Approach
How Multi-Agent Systems can replace most of Manual ML Validation decisions - The Karpathy Loop Approach

A fraud ring activates at 11 PM. Your detection model starts missing transactions it would have caught three months ago. By Wednesday morning, your monitoring dashboard is red. The challenger model is ready; it was trained last week, sitting in staging, waiting for the green light. It will not deploy until Friday. Not because the model or the infrastructure is not ready. This delay occurs because a data scientist must personally execute roughly twelve sequential quality gate decisions: schema validation, completeness checks, calibration comparisons, distribution parity tests, SHAP explainability reviews, threshold sensitivity analysis, and regulatory compliance sign-offs. Each one feeds the next. Each one requires a human to open a notebook, run cells, interpret output, and make a call. Rather than making scientific decisions, the data scientist is executing a deterministic checklist that was fully specifiable before they sat down. The fraud ring has a three-day window. This is not a technical failure. It is an organizational design failure wearing a technical pipeline as a costume. The Root Cause: Disguised Determinism Machine learning validation pipelines are frequently characterized by an over-reliance on manual expert intervention, a practice often misattributed to the inherent complexity of the domain. In practice, these validation gates are fundamentally deterministic: they involve executing computational functions and evaluating results against predefined scalar thresholds. Consequently, the human role in such instances is less about expert judgment and more about threshold enforcement. This operational bottleneck is often perpetuated by a reliance on ad hoc diagnostic notebooks and serial approval processes. These artifacts fail to distinguish between decisions requiring subjective cognitive assessment and those where constraints can be codified a priori. Within a well-instrumented validation pipeline, empirical evidence suggests that approximately 80% of decision gates depend solely on the clear definition of success criteria, leaving only 20% to require genuine human judgment. This realization establishes the foundational requirement for agentic validation architectures. The Karpathy Loop The iterative refinement principle of Andrej Karpathy is distilled as follows: an agent reads output, evaluates it against a scalar success metric, selects the next action from a defined action space, executes, commits or rolls back, and repeats. The agent needs three things: A clear objective metric, which is the scalar number that defines successA defined action space, which is the set of things it is allowed to tryMemory of what not to try again, which is a persistent record of dead ends The 5-Gate Architecture Four gates are autonomous, while one is human. Here is what they look like: Gate Function Outcome/Signal Gate 1: Data Quality Agent Validates feature dataset against schemas, completeness, and distribution. Fail signals Human Alert. Gate 2: Calibration Agent Executes calibration strategies (Platt, isotonic, etc.) against scalar constraints. Escalate to Human Alert on exhaustion. Gate 3: Distribution Parity Agent Compares production/challenger distributions; calculates KL-divergence. Fail signals Human Alert (Regulatory evidence). Gate 4: Explainability Agent Uses SHAP TreeExplainer for domain-sensibility/proxy detection. Flag for Mandatory Human Review if proxy detected. Gate 5: Human Approval Final review of package (results, logs, audit trail). Approve (Deploy) or Reject (Log). The do_not_try.md Memory Mechanism Every calibration strategy that fails the scalar constraint is written to do_not_try.md with the failure reason and observed metric value. On the next retraining cycle, the calibration agent reads this file before beginning exploration. There is no database or dashboard. There is no requirement to ask the senior data scientist who was there last time. Institutional memory is maintained as a markdown file. It is readable by any agent or human, survives personnel turnover, and accumulates across model families. This is the difference between an agent that is useful once and an agent that becomes smarter with each cycle. The Git Audit Trail Every decision in the Karpathy loop is a real git commit. The commit message is structured: ACTION strategy_name: metric_value versus threshold. Commit ID Action Description/Metric a3f91c2 KEEP rank_calibration: max_rate_delta=0.0% exp3 b7d44e1 DISCARD platt_scaling: 6.19% > 2.0% c91a3f0 BLOCKED isotonic_regression: do_not_try.md d02b5a8 PASS GATE3: histogram_overlap=94.2% >= 90.0% e445c17 PASS GATE1: 847291 rows, 0 null violations This git log is the compliance record. It is not a separate audit database, a dashboard someone maintains, or a PDF generated after the fact. Any reviewer can run git log --oneline and see every decision the agent made, in sequence, with the metric that justified it. It is immutable, append-only, and human-readable because it is produced as a byproduct of good engineering practice. Results From the POC Efficiency Metric Manual Pipeline Agent Pipeline Human Touch Points (per cycle) ~12 ~2 Gates Auto-Resolved 0% ≥80% Calibration Escalation Rate 100% ≤15% Retraining Cycle Time Days Hours Dead-end Re-exploration Frequent Eliminated The 83% reduction in human touchpoints is not from removing human judgment. It is from routing human judgment to the decisions that actually require it. The Recursive Nature of Architectural Development The efficacy of the proposed validation pipeline is derived from a recursive design philosophy; the architecture of the system mirrors the process by which it was constructed. Throughout the development phase, an AI-driven coding agent executed the design-build-test-debug lifecycle. In this configuration, the human researcher functioned as a director, defining high-level objectives and gate-specific success criteria, while delegating the implementation and validation logic to the agent. The conversation transcripts effectively served as a structured audit trail of decision-making, while the session context acted as a meta-level artifact, precluding redundant design iterations. This alignment, featuring the Human as Director, AI as Execution Engine, and explicit scalar metrics as success conditions, suggests that agentic workflows are not merely useful for model validation but are fundamentally transformative for the design process itself, functioning across varying levels of abstraction.

By Amey Farde
Building an AI Agent That Converts Production Failures Into Regression Tests
Building an AI Agent That Converts Production Failures Into Regression Tests

Production failures often contain enough evidence to explain what went wrong, but not enough structure to become an executable test. A trace may expose the failing request path, a log may contain the exception, and downstream spans may reveal the dependency response that triggered the defect. The useful engineering step is to transform that evidence into a deterministic regression test rather than another incident summary. Recent bug-reproduction systems follow the same principle that a useful reproducer should fail on the buggy revision for the reported reason and become passing evidence after the defect is fixed. Issue2Test and ReProAgent both use execution feedback instead of treating test generation as a single prompt-and-response operation. Start From the Incident Evidence The agent should begin from a machine-readable incident envelope, not a copied stack trace. OpenTelemetry’s stable log data model includes TraceId and SpanId, while its exception conventions associate exception records with the corresponding span context. W3C Trace Context standardizes traceparent for propagating trace identity across service boundaries. Those identifiers allow the failing execution path to be reconstructed without forwarding an entire observability dataset to a model. A small adapter can convert an alert into the minimum evidence required by the agent: Java FailureContext buildContext(Incident incident) { Trace trace = telemetry.getTrace(incident.traceId()); Span failed = trace.failedSpan(); return new FailureContext( failed.operation(), failed.exception(), trace.parentPath(failed), trace.downstreamCalls(failed), repository.revision(incident.deploymentId())); } The deployed revision is essential. A regression test generated against current source can target code that has already moved away from the production state. The incident should therefore resolve to the commit, image digest, or equivalent immutable revision that produced the telemetry. The trace supplies runtime evidence, and the repository supplies the code that interpreted it. Telemetry also requires reduction before model access. Request bodies, authorization headers, customer identifiers, and database values are rarely necessary to reproduce control flow. OpenTelemetry documents Collector processors to remove attributes, filter records, redact attributes, and transform values before export. Those controls should run before failure context reaches the agent rather than relying on a model to ignore sensitive fields. Reduce the Failure to Executable Context Raw traces are too broad for test generation. The agent needs a compact slice containing failing application frames, the request shape, relevant downstream interactions, and nearby tests that define local conventions. ReProAgent’s 2026 design separates bug localization, root-cause analysis, test planning, and test generation, combining repository retrieval with runtime interaction. Its results support treating reproduction as a staged, tool-using process rather than direct code completion. For a checkout failure, an error span may show InventoryClient.reserve() followed by a NullPointerException after the inventory service returned HTTP 503. Retrieval should locate InventoryClient, the calling checkout path, exception mapping, and existing checkout tests. Unrelated controllers, persistence code, and complete trace payloads add noise without strengthening the reproducer. The resulting agent input can be expressed as an explicit contract: Java TestRequest request = new TestRequest( context.failureFingerprint(), context.relevantSource(), context.relatedTests(), context.downstreamResponses(), "Generate one deterministic JUnit regression test. " + "Do not modify production code. Do not assert the observed bug as correct behavior." ); That final constraint is critical. A model can produce a test that asserts NullPointerException simply because production emitted it. Such a test would pass on the buggy implementation and preserve the defect. Bug-reproduction benchmarks instead use fail-to-pass behavior where the test fails on the pre-fix revision and passes after the correcting patch. Recent research on LLM repair validation also finds that passing executions can provide little bug-discriminating evidence, making differential validation important. Generate the Test Against the Intended Contract The oracle should come from repository evidence rather than model invention. Existing tests, API specifications, exception policies, sibling implementations, and documented response contracts can establish intended behavior. When those sources conflict, the candidate should remain unresolved instead of receiving a fabricated assertion. Consider a production failure where inventory returned 503 and checkout converted a missing response body into an internal NullPointerException. Existing endpoint tests may establish that unavailable dependencies map to a stable 503 response with an INVENTORY_UNAVAILABLE code. The generated regression test can encode that contract while reproducing the recorded dependency behavior: Java stubFor(post(urlEqualTo("/inventory/reservations")) .willReturn(aResponse() .withStatus(503) .withBody("{\"code\":\"overloaded\"}"))); mockMvc.perform(post("/orders") .contentType("application/json") .content(failureRequest)) .andExpect(status().isServiceUnavailable()) .andExpect(jsonPath("$.code").value("INVENTORY_UNAVAILABLE")); WireMock can match HTTP requests and return predefined responses, and it supports fixed or randomized delays and lower-level fault simulation. That allows a recorded external condition to become a deterministic test setup rather than a dependency on a live production service. Close the Loop With Execution Feedback Generation should be treated as the first candidate, not the final artifact. Issue2Test refines tests using compilation and runtime feedback, while ReProAgent includes runtime interaction throughout reproduction. A practical agent should compile and execute every candidate in an isolated checkout of the incident revision. Java TestCandidate refine(TestCandidate candidate, FailureContext context) { for (int attempt = 0; attempt < 4; attempt++) { TestRun run = sandbox.run(context.revision(), candidate); if (run.compiles() && reproduces(run, context)) return candidate; candidate = model.revise(candidate, run.diagnostics(), context); } return TestCandidate.rejected(); } The reproduces check should be stricter than “test failed.” It can verify that the expected application path was reached, the recorded downstream condition was exercised, and the observed exception or response fingerprint overlaps the incident. Compilation failures feed back into correction, a test that fails before reaching the target path is rejected and a test that passes on the buggy revision is not a reproducer. Once a fix exists, the same test should run against both revisions. ReProAgent defines fail-to-pass rate around exactly this distinction: failure on the buggy state and success after the issue-resolving patch. Differential execution is stronger evidence than asking a model whether generated code appears correct. Make the Test the Durable Artifact After deterministic replay, the reproducer can enter the normal test suite. JUnit treats failed assertions and uncaught exceptions as test failures, so ordinary CI can enforce the regression once the test is valid. Normal execution should require neither production telemetry nor another model call, and incident secrets should never be embedded in the generated fixture. A practical CI handoff can also preserve provenance without preserving raw incident data. A small metadata record can contain the incident identifier, source revision, generated test path, reproduction fingerprint, and validation command. That record makes regeneration and review easier while keeping the committed test independent of the observability backend. The test itself remains the executable source of truth. In practice, the generated test is verified under strict CI controls before ever reaching the main suite. The agent’s changes (adding the new test) occur on an isolated branch or worktree, and the CI pipeline runs git diff to confirm that only test files were created or modified, any application code changes cause an immediate failure. The test is then run against the original codebase to confirm it reproduces the production failure, and again against the patched build to ensure it now passes. Any anomaly (for example, the test accidentally passing on the buggy code or still failing after the fix) triggers a manual review. Meanwhile, any necessary fixtures from the incident (such as specific database records or request parameters) are set up in the test so it precisely mirrors the failure scenario. Metadata from the failure (stack trace, error message, etc.) is included in the commit or PR for traceability. This enforces that each generated test is precise and verifiable in CI before the developer ever sees it. Production observability becomes substantially more valuable when failures can be converted into executable evidence. The reliable pattern is to correlate telemetry to the deployed revision, reduce that evidence to the failing path, derive assertions from existing contracts, generate a deterministic test, and repeatedly execute it until the production failure is faithfully reproduced. The final acceptance criterion is demanding but clear: the test must fail for the real bug, pass after the real fix, and remain safe enough to run on every future change. That turns an AI debugging agent from a code generator into a controlled mechanism for converting operational failures into permanent regression protection.

By Uthej Mopathi DZone Core CORE

Monthly Top AI/ML Experts

expert thumbnail

Uthej Mopathi

Senior Software Engineer,
PayPal

expert thumbnail

Horatiu Dan

Senior R&D Software Engineer,
Tangoe

Horatiu is an R&D software engineer with 20+ years of experience in software development, mostly related to multi-tier enterprise applications. Throughout the years, as a certified Java and Spring Framework professional, he's been involved in all project lifecycle phases, from analysis, design and implementation to testing, maintenance and deployment of complex, high-impact products. The fields he contributed to address real-world business needs, in industries like Telecom and Maritime Transportation.
expert thumbnail

Pier-Jean MALANDRINO

CTO / AI Ambassador for the French Government, OSS maintainer,
SCUB

I am the Chief Technology Officer of a French digital services company, where I drive technology strategy, solution design, and R&D. I am also AI Ambassador for the French government's "Osez l'IA" (Dare AI) plan. My current engineering focus is low-bit LLM quantization. I built LLVQ, an independent from-scratch Rust implementation of Leech Lattice Vector Quantization, including a fused multi-shell CUDA decoding kernel and VRAM layouts for 2-bit weights. The work is published as a preprint (arXiv:2609.02652), with an open model on Hugging Face (Pier-Jean/Qwen3-4B-LLVQ-2bit) and an open-source repository (github.com/pjmalandrino/llvq). It also led to a merged upstream contribution to Hugging Face candle. I am the creator of Docling Studio (github.com/scub-france/Docling-Studio), an open-source visual inspection layer for document parsing, and I advise Karate Labs on UI product strategy and AI.
expert thumbnail

Pratik Prakash

Principal Solution Architect,
Capital One

Pratik, an experienced solution architect and passionate open-source advocate, combines hands-on engineering expertise with an extensive experience in multi-cloud and data science .Leading transformative initiatives across current and previous roles, he specializes in large-scale multi-cloud technology modernization. Pratik's leadership is highlighted by his proficiency in developing scalable serverless application ecosystems, implementing event-driven architecture, deploying AI-ML & NLP models, and crafting hybrid mobile apps. Notably, his strategic focus on an API-first approach drives digital transformation while embracing SaaS adoption to reshape technological landscapes.

The Latest AI/ML Topics

article thumbnail
AI on Top of a Dysfunctional System
Explore 10 Product Backlog anti-patterns that AI can make worse by masking dysfunction, missing evidence, and flawed decision-making.
October 2, 2026
by Stefan Wolpers DZone Core CORE
· 156 Views
article thumbnail
How a 30B Model Runs on Your Laptop
Big tech builds massive GPU data centers for LLMs, yet laptops can run models like Llama or Mistral locally. How is this done?
October 1, 2026
by Akash Lomas
· 462 Views
article thumbnail
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
The architectural divide between state-machine rigidity and agentic flexibility in financial systems, comparing stateful, multi-agent workflows.
October 1, 2026
by Soham Sengupta
· 409 Views
article thumbnail
Are Passphrases Still Secure in the Age of AI?
Passphrases have long been considered the most pragmatic answer to weak passwords, but the usage of AI raises a fair question: does "long and random" still hold up today?
October 1, 2026
by Constantin Kwiatkowski
· 402 Views
article thumbnail
Designing a Role-Aware Runtime for AI Avatar Systems
Use one shared runtime for AI avatars, with role profiles and adapters managing safety, tools, streaming, and each embodiment.
October 1, 2026
by Himanshu Gautam
· 375 Views
article thumbnail
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code
AI-generated code needs verifiable provenance linking intent, context, models, edits, approvals, commits, and artifacts across the software lifecycle.
September 30, 2026
by Uthej Mopathi DZone Core CORE
· 576 Views · 2 Likes
article thumbnail
Meta Wants to Run Your Business With AI — Microsoft and Salesforce Have a New Rival
Meta launches Enterprise Platform, bringing Muse, Business Agent and AI tools to companies as it takes on Microsoft and Salesforce.
September 30, 2026
by Ai Cerrudo
· 741 Views
article thumbnail
Engineering Self-Healing SQL Pipelines With LLMs: Validation, Guardrails, and Safe Recovery
Build self-healing SQL pipelines where LLMs propose repairs while deterministic validation, guardrails, and execution controls protect production systems.
September 30, 2026
by Uthej Mopathi DZone Core CORE
· 547 Views · 2 Likes
article thumbnail
OpenAI ‘o’ Leak: What We Know About ChatGPT’s Always-On Assistant Before DevDay
OpenAI may be testing an always-on ChatGPT assistant called “o,” with leaked references pointing to possible email functionality and ChatGPT Pro placement.
September 29, 2026
by DZone Staff
· 1,852 Views · 1 Like
article thumbnail
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A hands-on guide to Microsoft Foundry Document Intelligence SDK for extracting text, structured fields, and document data from PDFs for RAG and AI pipelines.
September 29, 2026
by Jubin Soni, FBCS DZone Core CORE
· 953 Views
article thumbnail
Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
In this article, you will learn why lines of code generated are almost meaningless as a measure of AI-assisted development.
September 29, 2026
by Gaurav Gaur DZone Core CORE
· 985 Views · 1 Like
article thumbnail
AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI coding is shifting from trusting models to constraining them with permissions, tools, and deterministic checks. BUBAS applies the same idea to business logic.
September 28, 2026
by Peter Verhas DZone Core CORE
· 810 Views
article thumbnail
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
Progressive disclosure replaces giant prompts with lightweight skill summaries and on-demand instructions, keeping AI agents focused and extensible.
September 28, 2026
by Akhil Madineni DZone Core CORE
· 899 Views · 2 Likes
article thumbnail
The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities
Reconcile AI plan changes with completed Temporal Activities using versioned plans, idempotent execution, and compensation rather than retroactive rollback.
September 28, 2026
by Akhil Madineni DZone Core CORE
· 755 Views · 2 Likes
article thumbnail
Why Incident Response Needs Memory, Not Just Intelligence
LLMs are excellent at reasoning, but effective incident response depends just as much on remembering previous incidents and organizational operational context.
September 28, 2026
by Akshay Pratinav
· 743 Views
article thumbnail
Software Quality Habits and AI
Software quality depends on daily habits. Learn how small habits compound and how AI can strengthen quality while weakening engineering judgment.
September 25, 2026
by Stelios Manioudakis DZone Core CORE
· 2,362 Views · 2 Likes
article thumbnail
Can Your Team Name the Work It Already Runs With AI?
The AI Workflow Inventory gives teams a shared list of recurring AI work, creating the foundation for A3 delegation decisions.
September 25, 2026
by Stefan Wolpers DZone Core CORE
· 1,190 Views · 2 Likes
article thumbnail
Locking Down the Enterprise: Data Security Patterns for AI Integrations
Learn how to secure enterprise data in AI systems using data classification, access controls, encryption, vendor controls, prompt security, and output validation.
September 25, 2026
by Balaji Venkatasubramaniyar DZone Core CORE
· 1,369 Views
article thumbnail
How Multi-Agent Systems can replace most of Manual ML Validation decisions - The Karpathy Loop Approach
This framework employs autonomous agents to run experiments against predefined metrics, determining success or failure without human input
September 24, 2026
by Amey Farde
· 1,708 Views
article thumbnail
Building an AI Agent That Converts Production Failures Into Regression Tests
AI agents can transform production telemetry into deterministic regression tests that reproduce failures and verify fixes automatically.
September 24, 2026
by Uthej Mopathi DZone Core CORE
· 1,507 Views · 2 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×