DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Methodologies

Agile, Waterfall, and Lean are just a few of the project-centric methodologies for software development that you'll find in this Zone. Whether your team is focused on goals like achieving greater speed, having well-defined project scopes, or using fewer resources, the approach you adopt will offer clear guidelines to help structure your team's work. In this Zone, you'll find resources on user stories, implementation examples, and more to help you decide which methodology is the best fit and apply it in your development practices.

icon
Latest Premium Content
Trend Report
Developer Experience
Developer Experience
Refcard #387
Getting Started With CI/CD Pipeline Security
Getting Started With CI/CD Pipeline Security
Refcard #399
Platform Engineering Essentials
Platform Engineering Essentials

DZone's Featured Methodologies Resources

AI on Top of a Dysfunctional System

AI on Top of a Dysfunctional System

By Stefan Wolpers DZone Core CORE
TL;DR: Polished Artifacts, Unchanged Decisions, Or Ten Backlog Anti-Patterns AI Makes Worse Add AI to a Product Backlog process that already struggles, and everything seems to improve within an afternoon. The problem is that polishing artifacts with AI doesn’t fix the root cause: the basis for the team’s decisions doesn’t change; AI only removes the visible discomfort that used to signal something was broken, along with some of the pressure to fix it. AI applied to a dysfunctional system – here, the Product Backlog process – makes the dysfunction look like progress. This is the first article of a new series that explains the uselessness of bolting AI onto a dysfunctional system. It explains why generative AI worsens ten Product Backlog anti-patterns, how polished AI artifacts can hide missing evidence and authority, and how teams can check whether their process is fit for AI. Thesis: Adding AI onto a dysfunctional system, for example, the Product Backlog, makes ten typical backlog anti-patterns worse, because polished AI artifacts hide missing customer evidence, decision authority, and feedback; the article explains these mechanisms, their costs, and how teams test whether their process is ready for AI. Disclaimer: I read Charniak/McDermott’s book on “Artificial Intelligence” decades ago; of course, I use AI for research, translations, proofreading, challenging story arcs and article structures, and summarization. It is a production tool, not a substitute for thinking. A Scenario You May Recognize Consider the following scenario: A Product Owner under pressure feeds stakeholders' requirement documents into an LLM, adds context such as the product roadmap or access to Jira. Soon, the Product Backlog holds 40 new items, each with a user story, five acceptance criteria, and an "estimated" size. Next Monday, Sprint Planning runs smoothly, and the stakeholders are delighted: finally, the team seems "ready." Now ask the four questions nobody asked: Which customer problem does this work address?Whose evidence says it matters?Which alternative did the team consider and reject?And who could have stopped it? If the answers are "whatever the stakeholders requested," "none," "none," and "nobody," the process didn't improve; it just stopped showing its flaws. (You may recall that Agile is really good at exposing problems, issues, and flaws.) What Must Work Before You Accelerate the Process Product Backlog management runs in a loop. What the team knows and assumes informs its product choices, and refinement exposes the uncertainty in those choices. Experiments, research, and delivered Increments then put those choices in front of customers and produce new evidence, which changes the next choices. That is why the Product Backlog is emergent, and every item comes with an implicit "based on what we know now." Work may legitimately enter the Product Backlog to investigate an open question; what matters is that people distinguish what they know from what they still need to learn and pick a suitable next step. If a team fails to feed valuable work into this loop, however, Scrum will be just as effective at building things no one needs: garbage in, garbage out. (Keep in mind that Scrum can be an effective delivery framework, but it completely lacks the product discovery part.) Refinement is a continuous activity within that loop, and it produces a shared understanding sufficient to make the next investment decision, including the decision not to build something at all. Selecting work for a Sprint happens later, in Sprint Planning, once the team has decided on a Sprint Goal: what is most valuable to create? A dysfunctional Product Backlog process usually announces itself: thin items, awkward Sprint Planning sessions, arguments about scope in the middle of a Sprint. That discomfort is useful, as it is the evidence a Scrum Master brings to the Retrospective and a Product Owner brings to stakeholders. AI can remove it without removing its cause. The Scrum Guide 2020 warning then applies literally: "Transparency enables inspection. Inspection without transparency is misleading and wasteful." Google's announcement of the 2025 DORA report, based on survey responses from nearly 5,000 technology professionals, describes the general pattern: "AI doesn't fix a team; it amplifies what's already there." The same announcement reports positive associations between AI adoption and both delivery throughput and product performance, alongside a continued negative association with delivery stability (Announcing the 2025 DORA Report). The findings cited do not establish the ten mechanisms described below; applying the amplification argument to them is my interpretation. AI can help repair a dysfunctional process if you deliberately use it to expose gaps: list the assumptions behind an item, find contradictions between a stakeholder request and the Product Goal (or the equivalent planning objective), or organize customer evidence the team already has. However, that only works while everyone stays aware of what the tool is doing. Polished artifacts fool people quickly into believing that AI is doing the real job, and teams need to be very careful not to fall into that trap. The useful question is whether an AI output makes evidence, assumptions, and alternatives easier to examine, or easier to skip. The dangerous intervention is the second: generating more apparently implementation-ready work before addressing the gaps. (AI, given the right context, is really good at papering over flaws and polishing its output.) What "Functional" Means Compared to a Dysfunctional System A functional Product Backlog process does not need to be flawless, fully documented, or unchanging. It needs to detect and correct poor decisions, which depends on three things: Direction: A customer problem or Product Goal makes a piece of work worth considering in the first place.Authority: Someone can challenge, change, or reject proposed work, and the organization respects that decision.Feedback: When delivery shows an assumption was wrong, later decisions change. The ten anti-patterns resulting from putting AI on top of a dysfunctional system below fail on one or more of these, and they fall into three groups: AI on Top of a Dysfunctional System, Group 1: Choosing Work Without Evidence 1. Prioritization by Proxy Stakeholders decide what goes into the Product Backlog, and the Product Owner passes their decisions on. With an LLM, a stakeholder now arrives with a ten-page requirements document, including personas and success metrics, produced in an afternoon. The polish does not create the power imbalance; it makes the existing one harder to confront, because rejecting a document that looks finished feels like obstruction and puts your career at risk. A better refinement session will not restore product management here, since the problem lies in who holds decision authority. The team becomes an internal development agency, only faster, and the question of whether something else would serve customers better never gets asked, because the polished document seems to have answered it already. 2. The Oracle Product Owner The Product Owner involves neither stakeholders nor subject matter experts, and the Product Backlog contains no research tasks, such as building prototypes or spikes. The AI version consults a model instead: synthetic personas and a generated market analysis. Source-backed desk research is legitimate work; the failure occurs when generated hypotheses acquire the status of observed customer demand. Product discovery gets skipped while looking complete, and the team then builds the wrong thing with considerable confidence, which may well be the most expensive way of building it, even if agentic coding makes building simpler. Without relevant customer evidence, generated personas and market narratives remain hypotheses, no matter how convincing they sound. 3. The Copy and Paste Product Owner The Product Owner breaks stakeholder requirement documents into smaller chunks. A model does this in seconds and returns neatly sliced items. Smaller items, however, do not make the underlying solution appropriate, nor does decomposing a document turn it into incremental learning. The Product Backlog commits prematurely to one solution, and nobody on the team has to understand the problem well enough to propose a cheaper or better-fitting one. The Product Backlog thereby fills with the stakeholder's solution, cut into Sprint-sized portions, and the later slices quietly turn into assumed commitments, which discourages the team from learning anything from the first slice before building the next one. 4. 100% in Advance The organization asks the team to rebuild a legacy application one-for-one, so the team prepares the complete Product Backlog upfront. With AI, a model reads the old codebase and produces an apparently complete specification with hundreds of items. The danger lies in treating that extracted description of the existing system as the approved specification for its replacement, including every workaround users have complained about for years. A rebuild is a rare opportunity to improve processes and usability; treating a generated inventory as fixed scope quietly closes that opportunity. I worked with a Scrum Team on such a rebuild for a large utility, and the team delivered two weeks early and under budget while streamlining processes and improving usability; all gains came from questioning inherited requirements. An extracted specification can even make existing behavior easier to question. AI on Top of a Dysfunctional System, Group 2: Turning Assumptions Into Apparent Readiness 5. Refinement Without the Team The Product Owner refines with the "lead engineer" and a designer, while the other Developers prefer creating code by herding agents. Asynchronous comments and preparation by a subset of the team are not dysfunctional as such; the problem arises when consequential disagreement never gets resolved. Estimation shows it best. The number was never the point; comparing estimates reveals whether Developers have the same work in mind before they start, as I argued in Estimates Are Useful, Just Ditch the Numbers. Accept an AI-suggested size without independent examination, and no one ever surfaces the split in opinion; the different ideas still exist, but nobody learns about them until the Sprint is underway. 6. Cosmetic Readiness All items look thoroughly detailed and estimated, yet the team has not applied INVEST in substance. AI makes every item look ready: correct template, complete fields, a checklist passed. Proper formatting, however, cannot prove what readiness depends on: credible evidence, understood uncertainty, workable dependencies, and shared understanding. A model can propose alternatives and analyze potential value; it cannot assert customer value or make an organization willing to negotiate. Sprint Planning then selects work that looks transparent but isn't, and the Developers discover missing understanding mid-Sprint, after implementation has started, and correcting it may require rework. 7. Acceptance Criteria Inflation The Product Owner covers every conceivable edge case without negotiating with the Developers. Ask a model for acceptance criteria, and you will typically receive a long list of plausible conditions. My heuristic is that three to five acceptance criteria usually suffice, and needing more often indicates the item needs splitting. The count, as such, is not the problem; the problem is criteria that expand scope without support or add constraints nobody examined, which the team's agents then build without evidence that real users need them. A separate risk arises when criteria start prescribing the solution instead of describing the required behavior: then the Product Owner drifts from the why and the what into the how, which belongs to the Developers, and the Developers end up negotiating with a document instead of a colleague. This is a particular challenge nowadays, when many organizations expect their product managers or product owners to be product builders, vibe-coding the prototype themselves. 8. The Last-Minute Product Backlog The team does not invest in Product Backlog management and rushes to fill the Product Backlog right before Sprint Planning. Before AI, the scramble was visible: thin items, a random assortment of stuff to fill the Sprint and appear busy, consequently resulting in a "Sprint Goal" nobody could explain. Now, an hour of prompting produces items that look as if the team had refined them for weeks. The preparation gap becomes much harder to see, although rework, confusion, and missed goals may still expose it later, by which point the cost has already been incurred. AI on Top of a Dysfunctional System, Group 3: Expanding Work Beyond the Capacity to Inspect and Sustain It 9. The Infinite Product Backlog Oversized Product Backlogs, Product Backlogs used as storage for ideas, and items nobody has touched for months predate AI. The economics change, though: generating candidates becomes nearly free, while assessing, ordering, and maintaining them still consumes the same human attention. If nobody moves anything to a separate list of permanently or temporarily discarded ideas, what I call the "Anti-Product Backlog," the most valuable items get buried in noise, and the ordering stops informing any decision. At that point, the Product Owner administers a database, and the team, facing hundreds of plausible items, stops reading the Product Backlog altogether. (Of course, you can instruct an agent to go through that list and flag all entries the agent believes offer no value. But this is a cleanup operation, not a proper process.) 10. What Technical Debt? The team ships feature after feature, and the Product Backlog reserves no capacity for bugs, refactoring, or platform maintenance. AI-assisted development increases the volume of change, and DORA's findings on delivery stability point to the risk when control systems, such as automated testing and fast feedback loops, are weak. This is, above all, a product decision: if the organization rewards feature output and never negotiates quality expectations or capacity trade-offs, AI will produce more features faster, and customers will experience the result as instability, not to mention that they are unlikely to pay for the flood of new features they never asked for. The Cost of Apparent Efficiency "We created the Product Backlog faster" is an incomplete business case, but typical for bolting AI onto a dysfunctional system. Consider an illustration, not a measured result: saving six hours of preparation has little value if unsupported scope then consumes three Sprints of delivery capacity, plus the review effort for AI output nobody inspected closely, the rework once users react, and the maintenance of features nobody needed. Meanwhile, the more valuable problem waits, and your competitors move ahead. When reporting after an AI rollout focuses on preparation time, the number of refined items, time spent on creating shared understanding, and the duration of Sprint Planning, none of these costs show up. The dashboard looks like a success while the expensive part of the path, building and maintaining the wrong things, sits in a different budget line and appears months later. Therefore, evaluate the whole path, from identifying a customer problem to learning whether the delivered change helped, including the effort it takes to inspect what the model produced. The costs also compound: every unexamined decision that "worked" teaches the organization that examination is optional. You can apply the A3 Delegation System to the problem to get a better understanding of where AI may support the process, where you should avoid using AI, what "good quality of AI outputs" looks like, and how to bring transparency to your team’s use of AI. Conclusion: Can Your Process Reject the Wrong Work? Before you accelerate product work by employing AI on top of a dysfunctional system, establish that your organization can question the work's justification, reject it, and learn whether it helped. Inspect a recent, important Product Backlog item and ask yourself which of the following criteria are true for your team: Grounding work in a problem: The team can name the evidence supporting the item and distinguish it from assumptions.Making choices: Someone can explain why the item took precedence and name a plausible request that was rejected or deferred.Exposing uncertainty: The people involved can name open questions and explain how they affect the next step.Exercising authority: The accountable person can change or stop the work when its justification fails.Learning from results: Feedback from delivered work has changed a later product decision. A first round producing convincing answers is rarely enough, so follow up and ask for recent examples: Which proposed item did the team stop or materially change because its evidence was weak?Which discovery or delivery result changed the next product decision?When did someone last exercise the authority the team claims to have? If nobody can point to such examples, generating more apparently ready work will only make the gap harder to see, although AI may still help you expose it. Key Questions This Article on the Effects of Employing AI on Top of a Dysfunctional System Answers How Does AI Make Product Backlog Anti-Patterns Worse? AI makes a dysfunctional Product Backlog process look functional, a typical first impression of bolting AI onto a dysfunctional system. It produces polished items, acceptance criteria, and estimates within minutes, while the basis for the decisions stays the same: missing customer evidence, unclear decision authority, and no working feedback loop. The visible discomfort that used to expose the problem, such as thin items or chaotic Sprint Planning sessions, disappears, and with it some of the pressure to fix the process. Should a Scrum Team Use AI to Write User Stories and Acceptance Criteria? Only if the team's Product Backlog process can already detect and correct poor decisions. AI generates user stories and long lists of acceptance criteria quickly, yet formatting cannot prove credible evidence, understood uncertainty, or shared understanding. Generated criteria also tend to expand scope without support; as a heuristic, three to five acceptance criteria usually suffice, and needing more often signals that the item needs splitting. How Can a Team Tell Whether Its Product Backlog Process Is Ready for AI? A process is ready when the organization can question the justification of work, reject it, and learn whether it helped. Test it with recent behavior instead of convincing answers: Which proposed item did the team stop or materially change because its evidence was weak? Which delivery result changed the next product decision? And when did someone last exercise the authority to stop work? Can AI Help Fix a Dysfunctional Scrum Process? Yes, when the team deliberately uses AI to expose gaps and stays aware of what the tool is doing. Useful applications include listing the assumptions behind a Product Backlog item, finding contradictions between a stakeholder request and the Product Goal, and organizing customer evidence the team already has. The dangerous use is generating more apparently implementation-ready work before those gaps are addressed. Why Does Creating a Product Backlog Faster With AI Not Save Money? Hours saved in preparation lose their value when unsupported scope consumes weeks of delivery capacity. The full cost includes reviewing AI output, reworking after users react, maintaining features nobody needed, and delaying work on more valuable problems. Reports that focus on preparation time, refined item counts, or Sprint Planning duration miss these costs because they appear months later in different budget lines. More
From Wild West to Context-Driven Engineering: How Culture Shapes Software Decisions

From Wild West to Context-Driven Engineering: How Culture Shapes Software Decisions

By Otavio Santana DZone Core CORE
Software organizations require rules to ensure predictability, knowledge sharing, and sound technical decisions. However, excessive rules, processes, and standards can obstruct progress if they eclipse desired outcomes. This longstanding tension in software engineering is now more pronounced as AI accelerates code generation, solution exploration, and implementation of change. The main challenge is not the number of rules, but how they fit with context, autonomy, and responsibility. Some organizations have minimal constraints, others tightly control decisions, and a few adapt principles to specific situations. Noticing these models helps engineering leaders assess their organization, understand associated risks, and move toward more effective technical governance. Software Architecture Is Also an Organizational Problem Software architecture is often defined by technical elements such as components, APIs, databases, infrastructure, quality attributes, and architectural styles. While these matter, they exist within a wider context. People design and maintain software, influenced by business constraints, communication structures, incentives, policies, and authority levels. Architecture is not only a technical artifact; it also reflects the environment in which teams make technical decisions. This broader perspective reflects the socio-technical nature of software engineering and is consistent with standards such as ISO/IEC/IEEE 42010, which recognize that architectural concerns encompass organizational, economic, regulatory, and community factors. Conway's Law shows that organizational communication structures often mirror system design. If organizational structure shapes software, then the distribution of authority, management of disagreement, response to failure, and governance of technical decisions likewise influence architecture. Therefore, engineering culture and governance are architectural concerns, not just management issues. Culture Shapes How Engineering Decisions Are Made If architecture is formed by its organizational context, organizational culture also directly influences engineers’ daily decisions. Sociologist Ron Westrum offers a system for understanding how organizations respond to problems and opportunities, particularly in high-risk environments. His model focuses on information flow, including cooperation, safe communication of bad news, responsibility management, failure investigation, and acceptance of new ideas. Westrum identified three main cultural patterns: pathological (power-oriented), bureaucratic (rule-oriented), and generative (performance-oriented) organizations. This distinction matters in software engineering, where architectural decisions depend on information moving across teams, technical boundaries, and leadership levels. Organizations that conceal problems, discourage disagreement, or fragment responsibility make different technical choices than those that discuss risks openly, encourage collaboration, and treat failures as chances to learn. DORA adopted Westrum's model in its research and consistently links generative, high-trust cultures to stronger software delivery and organizational performance. Culture isn't just the environment around engineering work; it shapes how teams interpret technical information, who can challenge decisions, how rules are enforced, and how architecture evolves. PathologicalBureaucraticGenerativePower orientedRule orientedPerformance orientedLow cooperationModest cooperationHigh cooperationMessengers “shot”Messengers neglectedMessengers trainedResponsibilities shirkedNarrow responsibilitiesRisks are sharedBridging discouragedBridging toleratedBridging encouragedFailure leads to scapegoatingFailure leads to justiceFailure leads to inquiryNovelty crushedNovelty leads to problemsNovelty implemented Mode 1: Wild West Engineering Wild West Engineering refers to environments with weak governance, unclear ownership, and few shared engineering principles. Teams may appear highly autonomous, but such autonomy lacks architectural direction, reliable feedback, and consistent accountability. Technical decisions are frequently reactive and localized, with tools and approaches selected for immediate needs or personal preference rather than the wider system context. The problem is not simply a lack of rules. Experienced teams can operate effectively with minimal rules if they share strong principles and understand the impact of their decisions. Wild West Engineering occurs when autonomy exists without sufficient shared context, accountability, or coordination. Over time, architectural knowledge becomes siloed, similar problems are solved inconsistently, and changes become riskier as no one fully understands the system. Westrum’s pathological culture is a relevant comparison: cooperation is limited, responsibilities are avoided, bad news is suppressed, and failures result in blame rather than inquiry. In software organizations, this often creates a hero culture where a few engineers become indispensable for managing disorder. AI can worsen this problem; when teams generate code, add dependencies, or make architectural decisions without shared guardrails, inconsistency and technical debt can spread quickly. Indicators Typical signs involve unclear ownership, multiple solutions for similar problems, undocumented architectural decisions, reliance on tribal knowledge, fear of changing legacy components, recurring blame after incidents, and dependence on a few key engineers. Another important sign is resignation: engineers stop proposing improvements because changing the system seems riskier than leaving existing problems unresolved. How to Move Forward Imposing heavy bureaucracy is not the solution to chaos. The first step is to establish minimum shared guardrails: clarify ownership, document key architectural decisions, define core engineering principles, set baseline expectations for security and operability, and create safer ways to discuss failures and technical debt. The goal is to support autonomy with enough structure to ensure local decisions conform to a coherent system. Mode 2: Bureaucratic Engineering Bureaucratic Engineering represents the opposite extreme, where governance is pervasive. Standards, approval processes, architectural reviews, mandatory technologies, and detailed procedures dictate how teams operate. While this structure creates predictability and consistency, problems arise when adherence to process outweighs knowing its purpose. Teams may focus on compliance rather than evaluating whether rules remain relevant. Westrum’s bureaucratic culture reflects similar traits: limited cooperation, narrow responsibilities, minimal cross-boundary collaboration, and a tendency to view novelty as problematic. In engineering, this occurs when architectural decisions are centralized, outdated standards persist, and exceptions are hard to secure. While these mechanisms may guarantee a baseline of quality, they may also limit performance. The core issue is not the presence of rules, but the erosion of judgment. Governance becomes rigid when organizations treat principles as permanent mandates and compliance as a substitute for engineering quality. AI can increase this tension, as organizations may respond to new risks with broad restrictions, approved-tool lists, or complex authorization processes. Although these measures may reduce risk, they can also hinder teams from pursuing valuable opportunities, making the organization safer but less nimble and flexible. Indicators Typical signs include rules missing explicit justification, standards defended only by “we have always done it this way,” centralized approval for minor technical decisions, difficulty obtaining exceptions, and architectural decisions that ignore local context. Status quo bias helps explain this persistence. Samuelson and Zeckhauser found that people often prefer existing or default options. In engineering organizations, this reinforces architectural inertia: once a technology, methodology, or rule becomes standard, replacing it requires far more justification than maintaining it. As a result, rules persist not because they are optimal, but because they are easier to retain. Organizational silence is another key warning sign. Morrison and Milliken define this as a collective tendency to withhold concerns when employees believe speaking up is unwise or ineffective. In engineering, this occurs when developers privately disagree with decisions but remain silent in meetings, believing that contesting established processes will not lead to change. Persistent silence should not be mistaken for agreement; it may signal that the organization discourages dissent. Another warning sign is when success is measured mainly by process compliance rather than software outcomes. Teams may meet all requirements yet become slower, less innovative, and disconnected from the initial intent of these controls. How to Move Forward To move beyond Bureaucratic Engineering, organizations should redefine governance rather than eliminate it. Leaders must frequently review the purpose of rules, distinguish primary constraints from preferences, and replace unnecessary approvals with clear principles. Teams need defined boundaries, authority to make decisions within them, and a clear process for exceptions when justified by context. A practical test is to ask whether a rule is still justified by its intended outcome. If no one can explain the risk it addresses, the value it protects, or evidence of its need, the organization may be maintaining the process for its own sake. The aim is to shift from prescriptive control to outcomes-based control, with accountability and feedback. Mode 3: Context-Driven Engineering Context-Driven Engineering integrates governance with local judgment. Rules are tools for managing risk and ensuring consistency, not ends in themselves. Teams work within defined principles, constraints, and responsibilities, while retaining the autonomy to adapt decisions to their particular technical and commercial context. The core assumption is that the same rule may yield different outcomes in different environments. This approach corresponds with Westrum’s generative culture, defined by high cooperation, shared risk, cross-boundary collaboration, inquiry following failures, and openness to innovation. In engineering, teams are expected to follow standards, understand their purpose, and recognize when exceptions are warranted. Architectural decisions are treated as explicit trade-offs, not automatic pattern applications. Teams evaluate technologies such as microservices, hexagonal architecture, event-driven systems, or specific cloud platforms based on the problem, constraints, and desired outcomes, rather than assuming they are universally correct. Applying a rule in the wrong context can lead to poor decisions. The same applies to architectural patterns: they address recurring problems in specific contexts, and applying them without considering context can create anti-patterns. For example, a lifeboat is essential in the ocean but irrelevant in a desert. Similarly, providing water is life-saving for someone dehydrated, but ineffective for someone drowning. A tool or rule's effectiveness depends on the context in which it is used. This environment requires greater maturity, not reduced governance. Autonomy must be balanced with accountability, feedback, and a readiness to revisit decisions as new information arises. AI fits well with this model, as its use can be governed based on context and risk: low-risk tasks may permit broad experimentation, while sensitive data, production changes, or high-impact decisions demand stricter controls. The objective is responsible judgment within clear boundaries, not unrestricted freedom. Indicators Key indicators include clear engineering principles, documented decision rationales, explicit ownership, and constructive disagreement. Teams should be able to explain both the rule and its rationale. Exceptions are allowed but must be justified. Failures prompt learning rather than blame, and architecture teams periodically review standards instead of letting them become outdated. Another strong indicator is that teams can contest established practices without equating disagreement with disloyalty or process violations. A context-driven organization distinguishes between non-negotiable constraints and situation-dependent choices. Security, legal requirements, data protection, and regulatory obligations remain strict, while implementation details such as architectural style, framework selection, or deployment strategy are evaluated locally. This distinction ensures autonomy does not devolve into unstructured or uncontrolled practices. How to Sustain This Model The main risk is regression. Insufficient discipline can lead to disorder, whereas excessive concern for consistency can create bureaucracy. Upkeeping this model requires ongoing review of principles, active feedback loops, documented architectural decisions, clear exception processes, and leaders willing to reexamine their assumptions. A practical test is whether the organization can answer four questions: What problem are we solving? Which constraints are truly non-negotiable? What trade-off are we accepting? What evidence would prompt us to revisit the decision? When teams can answer these questions consistently, governance functions as a learning system rather than a control mechanism. Conclusion Three Engineering Governance models inspired by Ron Westrum Engineering organizations improve not by simply adding or removing rules, but by ensuring teams understand their purpose, have the required context, and are trusted to exercise judgment within clear boundaries. Wild West Engineering lacks the structure for long-term growth, while Bureaucratic Engineering frequently leads to rigidity and limits learning. Context-Driven Engineering seeks balance through prioritizing principles over prescriptions, fostering autonomy with accountability, and encouraging decisions based on trade-offs rather than routine. As AI accelerates code generation and change, maintaining this harmony is increasingly important. The most adaptable organizations will foster responsible, context-aware decision-making. More
Can Your Team Name the Work It Already Runs With AI?
Can Your Team Name the Work It Already Runs With AI?
By Stefan Wolpers DZone Core CORE
Three Hidden Traps That Shape Software Engineering Decisions
Three Hidden Traps That Shape Software Engineering Decisions
By Otavio Santana DZone Core CORE
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures
By Sayan Chatterjee
Architecting Production AI Across Clouds: Patterns That Decide System Survival
Architecting Production AI Across Clouds: Patterns That Decide System Survival

Most enterprise AI post-mortems do not blame the model. They blame the storage tier that starved the accelerators, the identity policy that over-granted access, the cost model that ignored egress, the forecast that leaked future data, or the region that failed and took a business process with it. The hard part of production AI was never intelligence. It was the engineering discipline around it. This article distills the architectural patterns that decide whether a cloud AI system is trustworthy at scale, spanning infrastructure, identity, cost, operations, the applied domains, low-code assembly, platform selection, and multi-cloud resilience. It is written for engineers who have to keep these systems running, not for a keynote. Infrastructure: The Interconnect Is the Bottleneck Distributed training is a systems problem before it is a machine learning problem. When a job spans many graphics processing units (GPUs), the fabric connecting them (e.g., NVLink within a node, InfiniBand, or a vendor fabric across nodes) frequently caps throughput more than raw compute does. Accelerators wired through an ordinary network idle while they wait to synchronize gradients. Storage is the symmetric constraint. If the file system cannot deliver data at the rate the accelerators consume it, utilization collapses. The pattern is a tiered design: Hot tier: parallel or block storage feeding active training at high input/output operations per second (IOPS).Warm tier: recent data staged for quick promotion.Durable lake: object storage providing petabyte-scale durability, partitioned and lifecycle-managed underneath. Two cost drivers hide from the pricing page: data egress (moving data across regions or out of a provider) and idle warm capacity. Optimizing only the advertised compute line item guarantees a surprise on the invoice. Identity Is the Perimeter In a service-to-service AI architecture, the network perimeter is gone; identity is the boundary. A zero-trust posture, where every request authenticates and receives least privilege, contains the blast radius when a component is compromised. Across providers, identity federation is the load-bearing pattern: a principal authenticates once and is recognized everywhere, so access is granted and revoked centrally instead of reconciled across three identity systems. Policy must travel with the workload; a rule enforced on one cloud and forgotten on another is not a policy. Model authorization is the emerging frontier. As models call tools and take actions, the question moves from who can query this model to what may this model do on a user's behalf. Least privilege applied to an autonomous agent is the boundary between useful and unbounded. Cost and Operations Are a Control Loop Cost management is not a spreadsheet; it is automation. Consistent resource tagging across every cloud is the prerequisite for attribution. On top sit budgets, alerts, and automated remediation that throttles runaway spend before it escalates. Site reliability engineering (SRE) supplies measurable targets. For AI workloads, the golden signals extend beyond latency and errors to accelerator utilization, queue depth, and prediction quality. A model can be fully available and quietly wrong, so define a service level objective (SLO) for output quality, not just uptime. Three techniques earn their complexity: Spot or preemptible capacity plus checkpointing cuts training cost sharply when jobs resume cleanly after reclamation.Predictive scaling anticipates load instead of reacting to it.LLM inference optimization becomes architectural: batch requests, cache frequent responses, route easy queries to smaller models, reserve the expensive model for queries that need it. The Applied Domains Share a Spine, Differ in Physics Vision is byte-heavy. High-resolution images and video streams make the data and network layers dominant. For real-time video, decouple frame capture from analysis and sample frames rather than processing every one. Critically, a business-rule layer, never the model alone, owns consequential decisions. Every extraction should carry a confidence score used as a routing gate: Python def route_extraction(field, threshold=0.90): if field["confidence"] >= threshold: return "auto_process" return "human_review" Language is byte-light but semantically treacherous, and because it replies directly to users, errors are visible. The defining risk of generative systems is hallucination. The strongest architectural defense is retrieval grounding, forcing answers from verified sources with citations: Python def answer(question, knowledge_base): passages = knowledge_base.search(question, top_k=3) context = "\n".join(p.text for p in passages) prompt = f"Answer using ONLY this context.\n{context}\n\nQ: {question}" return model.generate(prompt), [p.source for p in passages] Forecasting is defined by time order. You cannot shuffle a time series into random splits, and the most common failure is data leakage, using information unavailable at prediction time. Test on a fair, time-ordered holdout, and always emit a prediction interval; a point forecast that hides its uncertainty invites overconfident decisions. No-Code and Low-Code: Governed or Ungoverned No-code and low-code platforms collapse build cost from a scoped project to an afternoon, which is why adoption is exploding. The symmetric risk is sprawl: hundreds of ungoverned flows handling sensitive data, owned by no one. Govern with guardrails, not gates. Restrict which connectors and data sources are permitted, assign an owner and an SLO to every production flow, then let builders move freely inside the boundary. The goal is to make the safe path the easy path. Platform Selection Without Self-Deception Vendors all claim to be fastest, cheapest, and most reliable. Benchmark to replace claims with evidence: Latency: report percentiles (p95, p99), never averages that hide the slow tail.Quality: measure on your own representative data, not a public leaderboard.Cost: model total cost of ownership, including transfer, storage, idle capacity, operations, and migration, not the headline compute rate.Reliability: verify the platform meets your recovery time objective (RTO) and recovery point objective (RPO). Combine dimensions in a weighted scorecard whose weights are fixed before scores are seen. Adjusting weights afterward to crown a favorite converts analysis into rationalization. Multi-Cloud Resilience: Design for the Day a Cloud Fails For systems a business cannot lose, a single provider is a gamble. Multi-cloud resilience deliberately places critical workloads so no single provider failure takes the business down, applied only where the cost of failure exceeds the cost of prevention. Predict rather than react. Combine leading signals into a health score and fail over proactively: Python def health_score(latency_ms, error_rate, saturation): latency_factor = max(0, 1 - (latency_ms / 1000)) error_factor = max(0, 1 - (error_rate / 0.05)) saturation_factor = max(0, 1 - saturation) return round(0.4*latency_factor + 0.4*error_factor + 0.2*saturation_factor, 3) Kubernetes makes workloads portable; data replication (with the consistency-versus-availability trade-off decided per workload) keeps data ready on the other side; and a portable foundation of federated identity, uniform policy, and centralized monitoring makes failover routine rather than heroic. The discipline that separates real resilience from a slide deck is rehearsing failure on purpose. An untested failover path is a promise, not a capability. The Judgment Layer Across every layer, value came not from the most powerful component but from the judgment applied to it: matching effort to problem difficulty, keeping humans on consequential decisions, measuring before deciding, building governance in early, and designing for change. Tools will churn; foundation models will make today's designs look quaint. That is precisely why principles outlast product knowledge. The scarce resource in enterprise AI was never intelligence. It was judgment, and judgment does not ship from the cloud.

By VenkataSrinivas Kantamneni
Alert Fatigue as a System Design Problem: Engineering On-Call Reliability in Modern SRE Teams
Alert Fatigue as a System Design Problem: Engineering On-Call Reliability in Modern SRE Teams

Once upon a time, site reliability engineering rested on a linear assumption: monitor more, detect early, and you’ll recover faster. The rise of alert fatigue makes modern SRE teams realize otherwise: Ramadass's (2025) paper, Building an AI-Powered Observability Pipeline for Modern System Reliability, cited research that discovered that: More than two-thirds (82%, actually) of institutions experience alert spikes constantly.Most traditional monitoring tools generate approximately 2,100 alerts daily, with about 70% of them unnecessary and safe to ignore.66% of SRE professionals stated that increased false alerts lead to fatigue, potentially causing them to miss serious issues. How Should We Describe This Situation? Vigilance or Noise? Collaborative systems such as SaaS, third-party APIs, and microservices enhance the degree of observability and notification within systems. Everything is monitored, and occasionally these dependencies may duplicate alerts. When systems request superhuman attention, on-call engineers become fatigued rather than lazy or sloppy. Instead of swift action, alerts are responded to with mistrust. Reliability vs. Experience vs. Metrics Traditional alerting metrics follow traditional reliability practices, that is, error rates, uptime percentages, latency, etc. Although these are essential, they are not actual mirrors of how operators or users experience reliability. Operators may expect reliable alerting to inform decisions, while users may simply define reliability as how well a system enables them to fulfill their intentions. If alerts do not clearly connect to the user experience, there is a gap between detection and action. Over time, the gaps lead to fatigue. On-call engineers begin to “reasonably” ignore these alerts. Why worry over alerts that are not logically related to user outcomes? They may assume. Over time, organizations may end up paying dearly for real issues because alerts were missed or delayed. An On-Call Engineer Experience Here is a typical example of a system design problem an on-call engineer or SRE team may face: 01:15 AM Alert: Latency spikes on a third-party API.01:16 AM Alert: Retry queues are filled.01:16 AM Alert: Timeout alert storms on three dependencies.01:17 AM Alert: Error-rate notification on unrelated endpoints.01:18 AM Alert: Memory and update alerts. And this sequence of alert storms continues, with the on-call engineer receiving more than 20 alerts in just four minutes. The system seems to pass standard observability SRE practice. But what about the long-run reliability suspicions that the bugging signals may create? In this case, the teams are not just grappling with response speed but also with the amplification of confusion when critical alerts are mixed with non-actionable ones. When Detection Outpaces Interpretation We can’t rule out the fact that monitoring in the past decades has taken an advanced leap. And we might be at its cloying stage, where system detection software is outpacing on-call engineers’ interpretation. Systems are “wonder-full” when it comes to identifying when something seems “off.” However, they rarely give explicit descriptions to aid SRE teams’ understanding. An alert can indicate that a queue has exceeded its depth, but may not categorically state whether the issue is temporary or actionable, or whether users are affected. This occurrence spans dozens of dependencies, each with its own signal. The on-call engineer is kept puzzled about the best action to take at the right time. Hence, a reliable response could be excessive caution or delay as the engineer seeks to clarify the situation. The users are negatively impacted. Although the system met technical observability SRE standards, it failed operationally due to its opacity. The Hidden Cost of Alert Overload We rarely see the outcome of alert fatigue overnight. Its effects build up. Delayed response time accumulates. The aftermath incident review loses credibility. Engineers are skeptical of alerts and hesitate to decide first whether they are real or false. The cultural cost of alert fatigue is that on-call roles become a burden SRE teams endure rather than enjoy with a sense of responsibility. In the long run, engineers may feel they have no control over issues due to the confusion that multiple alerts create. Ironically, the same reliability problems that alerts were designed to solve are what they quietly create. Are Alerts Creating a False Sense of Safety? Lots of alerts may seem like a good thing or a sign of strong monitoring at first glance. But here is the truth: alerts could be hiding actual risk. As every deviation is notified, critical and minor alerts blend in. Teams begin to feel alert fatigue and delay response. Then, real problems begin to breed behind the scenes. Remember how SLAs could paint an illusory picture of safety? Similarly, alert volume could do so. Therefore, your SRE team should bind these caveats as the core of their modus operandi. Alerts shouldn’t replace action.Alerts shouldn’t be unsorted (by machines or humans).Alerts shouldn't be discarded. Alerts are signs that our systems need attention, and we should never be tired of listening. SRE Teams Designing Systems that Alert Smartly High-quality systems respond efficiently when dependencies fail. Instead of creating panic, they automatically degrade. SRE teams could design circuit breakers that could inhibit alert storms before they explode. They could also install bulkheads to prevent a single failure from spreading. There could be alert limits and a summary of conditions that resolve the problem of spamming. Instead of relying on metrics, system engineers could set up composite alerts that describe system states. For instance, it’s clearer if a system alert indicates, “Checkout degraded because of latency in payment dependency.” This composite alert is better than 7 alerts that say “Checkout Timeout.” The former shows impact, cause, scope, and urgency. Clarity clears fatigue. Noise does the opposite. Redesigning SRE: Human Reliability That Quells Alert Fatigue We have seen that technical designs may be great, yet other aspects of SRE remain wanting. One such area that could resolve a system design problem is humaneness. To avoid alert fatigue, our design choices must acknowledge human limitations. Therefore, we should accept that some alerts may not require immediate response. Conversely, not every anomaly should trigger an alarm. Understood silence could sometimes be a golden sign that nothing critical is wrong. Advanced SRE teams do not focus on events (or every deviation) but on the states of the system or infrastructure. They are guided by the question: What conditions really impact users, business objectives, or the system's overall health? To achieve this, engineers need to balance product understanding with technical operations. Then they can give a human touch to their designs. Designing systems for human reliability requires a high level of discipline. Site reliability engineers have to continually review, refine, and repair alerts and their trigger commands. Systems are like living organisms that need constant feeding of updates. The evolving nature of alerts could make a helpful one-time alert redundant or harmful in six months. On-Call as a Reliability Interface of SRE No doubt, humans have a role to play in ensuring reliability, but system designs that depend on heroic actions are built not with resilience but with fragility. Reliability is truly achieved when on-call engineers are guided by predefined scripts, models, runbooks, signals, and interfaces. These reduce the tendency to resort to fallible improvisations when issues arise. On-call engineers often take the appellation of “last point of call.” A careful look at their roles shows that they are intermediaries among complex systems, user experience, and consequences. We can thus see that the role of on-call engineers extends beyond problem resolution to stewardship. Conclusion Alert fatigue is a design problem. It often arises when on-call engineers prioritize detection over interpretation, or technical workability over user experience. The dependencies of modern SRE teams make it necessary to align technical alerts with human capability. Alert storms could wear out hardworking engineers who need to take a break. So, system designs need to account for human limitations, recognize that runbooks are better than on-the-spot improvisation, and prioritize clarity over opacity. Designs that account for these factors reduce or eliminate fatigue and preserve the very essence of alerts. In summary, reliability goes beyond resolving many problems to responding to what matters most. When teams can always trust their alerts, they will be more likely to follow up on new cases.

By Oreoluwa Omoike
Reliability Without Control: Operating SRE Practices in Platform–SaaS and API-Dependent Systems
Reliability Without Control: Operating SRE Practices in Platform–SaaS and API-Dependent Systems

Originally, back-end and front-end Site Reliability Engineering (SRE) were owned by teams. They code the programs, set up databases and infrastructure, and quickly spring to action at the beep of any anomaly. The advent of code vs no-code infrastructure, SaaS, API dependencies, third parties, and other modern systems seems to be eroding this authority. Mainstream and underdog companies now often leverage the significant advantages of outsourcing, collaboration, or delegation, which are usually accompanied by a silent clause: no or partial control. Unlike in previous systems, modern production is largely assembled rather than built from scratch. For example, a conventional SaaS product is built on interdependencies among payment processors, outsourced data infrastructure such as Amazon Web Services (AWS), messaging services, web hosting, design, AI inference APIs, authentication providers like Google, and more. These useful platforms and products are essentially outside teams' control stations, even though they critically impact users' experience. When they function effectively, you share the glory with the platforms. But when there is a system blackout, your users put you on your toes, even though you have no direct access to resolve the problem on time. Therefore, we shall be exposing SRE practices in platform-SaaS and API-dependent systems and how reliability is getting beyond the control of engineering teams and companies. Why Classical SRE Practices May Fail One major downside of SaaS and dependency on external platforms is that reliability control is often assumed to be in a team's hands, whereas it has been bargained. However, teams must reckon with the fact that the case is reversing. For example, traditional SRE models once alleged that: Service Level Indicators (SLIs) focus on availability or internal uptime and latency.Error budgets arise from changes teams make or deploy.Runbooks still suggest that teams can immediately reconfigure or directly work on faulty components. All these are becoming past cases, especially in platform-SaaS systems. You can have a system indicating 99.99% or even 100% uptime on the back end, while new users are struggling to sign up, probably because an authenticator provider is not fully functional. Dashboards and control panels may indicate green, but in reality, third-party payment APIs have been degraded. A New Definition of Reliability in Operating SRE Practices To resolve the new problem in site reliability engineering (SRE), there needs to be a conceptual shift from component health to an integrated, continuous user experience. Therefore, teams need to undergo a paradigm shift away from questions such as "Is our CPU working maximally?" “Is our API up?” “What are the error rates?” Instead, we should inquire: “Are users checking out seamlessly?” “How fast can they authenticate?” “Can they use the SaaS product to perform its key function?” These types of outcome-based questions span interdependent platforms beyond your full control. The login SLI needs to work with the identity provider; otherwise, its output is meaningless. If the checkout SLO skips payment authorization, then it's both fishy and unreliable. True, there may be some internal errors in a reliable system, but what really matters is an integrated multiplatform experience that the user enjoys. Error Budgets? An SRE Practice to Revisit How many teams would love error budgets to disappear when they give up control? But that’s not so. Instead, they are molecularized. When components of your systems are outsourced, the error budget doesn’t just fade away; it is instead transferred to the interdependent platforms. So, it’s better to plan for the fact that SaaS and API providers will consume some of your reliability budget. Doing so keeps you a few steps ahead and protects your business in the long run. Reliable SRE teams make decisions such as allocating part of their error budget to certain dependencies, setting acceptable parameters for degradation, and defining specific steps to take when a dependency exceeds the stipulated budgets. Here’s an example you can adapt: “We will accept payment authorization failure of 0.0% to 0.2% if it is caused by dependency instability. If it goes above that, we will turn on delayed capture or turn off promotions.” This SRE approach keeps you ready for downtime, as your systems automatically switch to planned or budgeted actions rather than relying solely on integrated platforms. What to Do When Failures Beyond Your Control Arise Actually, some failures may seem beyond your control. The more you attempt to resolve them, the more amplified they become. At this point, your team must adapt to the savvy absorption of such situations. Instead of focusing solely on retrial in an SRE approach, your team needs to design its processes and platforms. This could include failing selectively through circuit breakers, failing fast with timeouts, or failing visibly by keeping users informed. Some core settings should always remain non-negotiable and on standby. These could include the following: Read-only modes/cachesBulkheads that prevent a failure avalanche.Automated circuit breakersDeferred processing These reliable practices ensure there is some form of controlled uptime even when operations seem interrupted. Laser Observability That Proves Reliability In traditional SRE observability, the service boundary is usually the ultimate, but in most modern integrated SaaS platforms, this could be insufficient or worse, dangerous. Operators need to be aware of the actual dependency that is failing, how it is failing (e.g., errors or throttling), and how the failure affects the user experience. Accurate observability for platform-SaaS and API-dependent systems requires these four provisions: Specific dashboard and internal metrics for each vendor.SLI monitoring at the dependency level.Parallel tracing of all outbound calls.Simulation of real-time user experience and workflows. Essentially, whenever there is an emergency, operators should be able to promptly identify whether the source is internal or external. Accuracy and clarity facilitate swift response. Responding to Incidents Without Ownership Another distinct characteristic of modern SRE practice in platform-SaaS is how incidents are responded to. Without ownership, you often cannot debug on your own, roll back a bad deploy, or directly manage other issues. However, you can choose how your system responds by identifying when certain features are disabled, when signals to activate degraded modes are sent, when high traffic is redirected or shed, or when to notify users. To maintain reliability, incident response relies on runbooks to inform decisions. The following questions could help convert the technicality of runbooks to practical solutions: What is the impact on the customer?In what ways can we respond harmlessly?What can we reverse?What should we communicate externally? These questions help resolve incidents, mitigate losses, and intertwine reliability with sound judgment. Is Safety an Illusion in SLAs? SLA providers often readily contract for financial compensation when losses arise, but seldom give absolute reliability guarantees. You may not always expect vendors to consistently meet your availability goals or resolve an avalanche of outages. Safety is a critical consideration when building systems, because when users lose trust in a brand, compensation may not be able to redeem it. Therefore, advanced teams do not consider SLAs as safety nets but as risk pricing. They understand that contractual credits cannot replace trust, brand image, and some almost irredeemable damages. Human Factors in Platform-SaaS and API-Dependent Systems Dependency failures often escalate when cognitive load increases. There could be degraded performance, timeouts without error indicators, partial success, or inconsistent system behavior. Operators may not only focus on machines when dashboards lag or seem to lie. They examine the logs, failure history, or commands. Teams have to design systems with overrides and predictable degradation paths, and observability tools are beyond the failure systems. Reliability goes beyond the correct function of software; it's also about human operations. How Your SaaS and API Platforms Can Imbibe “Good” SRE Practice Effective SRE practices are modern. The following attributes know saas products and API-dependent platforms: Acknowledgment of lack of control very early.Ensuring reliability is embedded in the design.Measuring the outcomes of each SRE criterion or target, instead of just the components.Giving priority to clarity instead of trying to model or control everything because you do not own all the components.Making engineering and operations decisions and products as an integrated whole.Preparing for degradations as inevitable procedures when things fail. Your systems can be reliable if you anticipate failure and accept the reality. Conclusion Modern platform-as-a-service (SaaS) operates in a reliability-without-control manner, leading solid SRE teams to accept that they need to adapt when failures occur. It's simple logic: if you don't absolutely own everything end-to-end, then prepare for the worst: each dependency might fail. It's all about keeping the trust of your users and protecting your brand image.

By Oreoluwa Omoike
From Agile to the Product Operating Model
From Agile to the Product Operating Model

TL; DR: The “Agile to the Product Operating Model” Survey Results Between August 2 and August 10, 2026, 48 practitioners participated in my Agile to Product Operating Model (POM) survey, which tries to shed light on what is actually changing. Let me summarize the answers for you: the reported transformations change decision-making less than the Cagan framework suggests. Where respondents report improvements, they appear in delivery and collaboration rather than in business results. Unfortunately, the human side of the transition is the least encouraging part of the answers. Thesis: Product operating model transformations mostly change vocabulary and organizational structure while leaving the decision system — who decides what gets built, on what evidence, at what speed- largely untouched. AI is changing product decisions independently of POM transformations. Who Answered, and Why I Report Counts, Not Percentages Of the 48 respondents of the Agile to Product Operating Model, 26 work in organizations that have adopted a product operating model or are moving toward it (we refer to them as the “movers” from here on). The other 22 work in organizations that are not moving; these organizations are still discussing the issue, have decided against it, or have never considered it. So you get absolute counts, and every finding below is directional, not definitive. Additional limitations are: 24 of the 48 respondents are Scrum Masters or Agile Coaches, the roles under the most pressure from this shift.38 of 48 work in organizations with more than 250 people; the sample contains 2 respondents from startups and 2 from scale-ups. Apparently, the people who have participated in the survey are those who still care enough about agile practice to read a newsletter about it; people who left “Agile” entirely are structurally missing. Also, as I did not ask respondents to identify employers, the 48 participants report on their organizations, not 48 distinct verified organizations. Putting it all together, this is a report from inside large organizations, where the product operating model makes its boldest promises. One Fundamental Change in 26 Answers Asked which statement best reflects their experience so far, the 26 movers split like this: 9 say mostly relabeling, the same way of working with new vocabulary9 say real change in some areas, relabeling in others7 say it is too early to tell. Exactly one participant reports a fundamental change in how the organization decides what to build, which is the most interesting learning from the responses: changing the vocabulary, changing the organizational structure, and changing how decisions actually get made are three different things. Better Delivery, Inconclusive Business Results, Worse Morale The Agile to Product Operating Model survey asked movers to compare six outcomes against their previous way of working. Aggregating “somewhat” and “clearly” in each direction: Speed of delivery: Better (9), No change (10), Worse (1), Too early / cannot tell (6)Value delivered to customers: Better (8), No change (8), Worse (2), Too early / cannot tell (8)Business results: Better (3), No change (8), Worse (4), Too early / cannot tell (11)Team motivation and morale: Better (7), No change (4), Worse (9), Too early / cannot tell (6)Developers’ satisfaction: Better (4), No change (8), Worse (7), Too early / cannot tell (7)Collaboration with stakeholders: Better (9), No change (6), Worse (4), Too early / cannot tell (7) The table does not say that the transitions are failing: Delivery speed leans positive: 9 better against 1 worse.Customer value leans positive: 8 against 2.Collaboration with stakeholders leans positive: 9 against 4.Business results are inconclusive, with 11 of 26 unable to say and the rest split. The people dimensions are the only ones that lean consistently negative: morale at 9 orgs is worse than at 7, where it improved; developers’ satisfaction is similar: worse at 7 orgs and better at 4 orgs. Two ways to interpret this information, and the survey does not help to choose: (1) The improvements concentrate on the things enterprises already know how to optimize (flow, coordination, delivery), while the transformations may be falling short of the things the model claims to change most fundamentally. (2) Operational effects precede commercial effects, because business outcomes have longer feedback cycles, and 11 of 26 respondents explicitly cannot yet say. Consider that flow, coordination, or delivery are easier to measure and attribute than commercial effects, which is a significantly fuzzier area. Therefore, the interesting question is how this changes over time, not how it looks at one moment. The range of individual experience behind those aggregates is wide. One respondent, 12 months into adoption at a large financial-services organization, described the hardest part as “surviving the political game that the POM shift has been” and reports worse answers on five of the six dimensions. The other two respondents who have been at 12 or more months report the opposite: better value, better morale, and better collaboration. Three long-tenured accounts cannot agree on whether the practical transition mechanics, culture, affected products and services, or the organization itself makes the difference. (That is the curse of a small sample size.) Agile to Product Operating Model: The Empowerment Gap, and What Settles Uncertainty Cagan’s own test separates empowered product teams from feature teams: in a product team, the team is tasked with solving a problem and owns the solution. In a feature team, “the value and business viability are the responsibility of the stakeholder or executive that requested the feature.” The survey asked movers which pattern operates today, in the middle of their transitions: 9 of 26 report the empowered pattern (leadership sets goals and problems to solve, teams decide what to build)8 of 26 report that leadership decides which features get built and teams implement them.4 of 26 report stakeholder requests are filling the product backlog reactively.4 of 26 say it is genuinely unclear or contested right nowA single one reports teams setting their own direction. So even among respondents whose organizations are actively adopting a model whose entire premise is empowered teams, the “feature-list or reactive-backlog” pattern (12) outnumbers the “empowered” pattern (9). A related question asked what settles it when the organization is uncertain whether something is worth building: 6 said debate and prioritization before anything gets built5 said whoever has the most authority decides5 said things just get built and shipped, and an explicit decision rarely follows. 4 said research evidence5 said it varies too much to say. In total, only five respondents fall into the categories I pre-registered as evidence-led: research evidence, disposable prototypes, or ship-and-measure. There is one case among all the answers: a respondent reports that a disposable AI-assisted prototype reduces build uncertainty: build to learn, decide, and throw it away. That same respondent, 12 months into adoption, also reports better responses across all six outcome dimensions listed above. Of course, one case is not evidence of a relationship. It is a question worth asking in a larger sample, not a claim I am entitled to make here. In Little Code, Big Waste, I argued that cheap AI code removes the cost gate that used to force a should-we-build-this decision, and that “when generating plausible code becomes cheap, every hour spent building the wrong thing becomes waste that can now be produced at scale”. This survey says the prototype-to-decide pattern barely exists yet in this sample. Agile to Product Operating Model Survey Shows: The Two Transitions Run Separately If you believe the keynote version of events, AI is forcing organizations to redesign their operating model around it. The 26 movers report something different. Only one respondent says AI is a core reason for the change; 6 of 26 call it one factor among several; 16 say AI adoption runs in parallel, but separately from, the operating model change; and 3 report little or no role. Now the other side. Among the 22 respondents in non-moving organizations, 15 report that AI is changing how product decisions are made anyway: 6 noticeably, 9 in pockets, without any formal model change. Despite the small sample size, the data support a narrow claim: for most movers in this sample, AI and operating-model transformation are separate initiatives, while for most respondents in non-moving organizations, AI is changing product decisions without any model change at all. Formal operating-model change is neither a prerequisite for AI-driven changes to product decision-making nor, so far, organized around them. The survey did not measure how AI enters these organizations or who authorizes it, so I will not claim more than that. The disconnect between the two is already the finding. My Pre-Survey Hypothesis Scoreboard Here are my verdicts on my four hypotheses, H1-H4, that led to the creation of the Agile to Product Operating Model survey: H1, consultant-driven transitions produce more relabeling than leadership-driven ones: Relatively consistent, but no verdict: 3 of the 4 respondents whose transition is driven by consultants or a transformation office report “mostly relabeling,” against 4 of 16 in product-led or executive-led transitions. But those are just four participants. H2, the empowerment gap: The feature-list plus reactive-backlog patterns (12) outnumber the empowered pattern (9). That is observed in this sample, however, not established beyond it. H3, where AI drives the move, product judgment is the scarcest capability: incorrect on my side: Among the 7 respondents whose organizations treat AI as a core reason or a factor in the move, the most frequent scarcity answer is “delivery capacity is still the bottleneck” (3), ahead of product judgment (2). Across all 26 movers, the scarcity question splits three ways: “stakeholder alignment and decision speed” (7), delivery capacity (6), and product judgment (6). One distinction keeps H3’s rejection from closing the question. The survey measures perceived constraints, and perception rarely reflects reality accurately. AI may have changed the economics of implementation faster than organizations have updated their sense of where the workflow bottleneck sits. Whether delivery capacity objectively remains the constraint is a different question, and this survey cannot answer it. What it can say: practitioners today experience decision speed, delivery, and judgment as roughly competing constraints, and the judgment-scarcity era I anticipated is not the world they report living in. H4, organizations that build without explicit decisions report worse customer value than evidence-led ones: rejected as stated: 1 of 5 build-without-deciding respondents reports better customer value, against 2 of 5 evidence-led ones. Again, the number of replies is too small. Changing the Structure Without Changing the Decision System Three of the findings above belong side by side. Only 1 of 26 movers reports a fundamental change in how build decisions are made. Only 9 of 26 report the empowered decision pattern operating today. Only 5 of 26 fall into the evidence-led categories for resolving build uncertainty. Together, they suggest a more precise diagnosis than “the transformation is theater.” Transformations change three different layers, and the layers move at different speeds: Vocabulary changes first: Product operating model, empowered teams, outcomes over outputs; you get the idea.Structure changes second, and often really does change roles, reporting lines, team topologies, or artifacts.The decision system changes last, if at all: Who decides, based on what evidence, under what uncertainty, with what authority, at what speed. This survey reads like a snapshot of organizations renaming the first layer, reorganizing the second, and leaving the third largely untouched. That is why the 1-of-26 count deserves the weight I put on it earlier. One of the product operating model’s defining promises is a different way of deciding what to build. If only one of the 26 respondents experiences a fundamental change in that mechanism, the question is no longer whether the transformation is proceeding fast enough, but what is being transformed. That mental model offers one possible explanation for the outcome table: structural change could improve coordination and flow before it alters the quality of product decisions. Whether that explains the pattern here is impossible to establish from 26 responses. A long-term study of organizations and involved practitioners would need to test whether decision-system change predicts eventual business outcomes. (Consider that a hypothesis this survey generated, not a result it delivered.) Agile to Product Operating Model: Why Product Washing Is Easier Than Empowerment The survey cannot tell us why the decision system resists change. What follows is my hypothesis, argued from two decades of watching transformations, not a survey result. The tempting explanation is that these organizations are implementing the product operating model badly, and that a proper implementation would deliver. I spent those two decades watching the Agile community run exactly that defense. Every failed adoption was “not real Scrum.” The argument is unfalsifiable, and it taught an entire industry to blame practitioners instead of examining incentives. I will not run the same defense for the product operating model. In November 2024, I described Product Washing: the hollow adoption of product practices that “leaves companies stuck in the same old dynamics but with a new vocabulary,” transformation by reprinting business cards. My hypothesis for the mechanism: product washing is not an implementation failure but what enterprise incentives produce when you ask powerful people to redistribute their own power. The model demands that stakeholders with budget authority hand problem-selection to product leadership and solution-selection to teams. Budget authority is power, and in many large organizations, “product leadership” has become a new title for the same stakeholders who control the money. Meanwhile, middle layers face asymmetric career payoffs: a visible failure damages a career far more than a shared success advances it. Under those payoffs, routing decisions through committees and sign-offs is rational self-protection, and it does not vanish because the org chart was redrawn. Two honest caveats bound this hypothesis: First, in regulated industries, some of that scrutiny is very relevant: when a named person must answer to a regulator, a sign-off chain is accountability, and the respondents from banking and the public sector live with constraints no product operating model erases. The skill worth having is telling required governance apart from accountability theater; most organizations run both and label neither. Second, this survey is a cross-section, not a time series, and two competing explanations fit the same data: The transition hypothesis: decision authority changes slowest of all the layers, and these organizations are simply not there yet.The attractor hypothesis: enterprise incentives pull transformations toward renamed feature factories and hold them there. My incentive argument predicts the attractor. With only 3 respondents at 12 or more months, this survey cannot distinguish between them. That is a testable question for a future survey. The Non-Movers’ Catch-22 The 22 respondents in non-moving organizations deserve more attention than transformation literature usually grants them. Their top reasons for staying put: Leadership sees no need or has other priorities (15),Lacking the product-management maturity to build on (13), andThe cost and fatigue of yet another transformation (10). Only 6 claim their current way of working performs well enough. The dominant non-mover position is not a confident endorsement of the status quo. And inside the reasons sits a genuine Catch-22. Organizations supposedly need the product operating model because their product capabilities are weak, while 13 of 22 respondents say weak product capabilities are precisely why their organization cannot adopt it. A transformation that requires the maturity it promises to create is a hard sell to people who have already survived several that made the same offer. Conclusion: Three Questions for Monday Morning Three questions locate your own organization on this map: First: who decided the last thing your team built, and would your CPO name the same person? If the answers differ, you have found the gap between the model on the slides and the model in operation. Second: what settled your organization’s last genuinely contested build decision: evidence, debate, or seniority? “Whoever has the most authority decides” got 5 votes out of 26 in this survey. Be honest about whether your organization would add a sixth. Third: what would have to be true for a disposable prototype to settle the next contested decision instead? That is a political question, not a technical one: who would have to accept evidence as a tiebreaker, and what would it cost them? Forty-eight answers later, the better question is no longer which operating model organizations adopt. It is how they make decisions when the cost of trying something has collapsed while the cost of deciding has not: Who decides on what evidence, and how quickly can the organization act on what it learns? That is where my work is heading next.

By Stefan Wolpers DZone Core CORE
Code Generation Is Solved; Trust Is the Bottleneck
Code Generation Is Solved; Trust Is the Bottleneck

You have a checkout flow. You have 40 tests. They're green. Now: what happens when a payment webhook arrives after the user cancels? What happens when a retry lands on a session that already expired? What happens on the fourth failed attempt when autoRenew is off and the period boundary has already passed? You don't know. Not because you're careless — because a state machine with 6 states, 7 actions, and 3 payload values has thousands of reachable (state, action, data) combinations, and your 40 tests visit 40 of them. The bugs that page you at 2 am live in the other several thousand. Polygraph is a Claude Code plugin and standalone CLI that walks all of them. Why You'd Bother Polygraph is for stateful code: reducers, workflow engines, protocol handlers, session managers, order state machines, anything with a dispatch(state, action) shape. If your code is a pile of pure functions, go use property-based testing. If it's a state machine, keep reading. Narrow on shape, not on language. The model reads your source in whatever it's written in — you name it once as lang in the contract — and the trace format is just NDJSON, so any runtime that can log a {pre, action, data, post} line per step can feed it. What's always JavaScript is the derived spec and your rules, because those are what the replayer and model checker execute on Node. What you get back is not a lint warning. It's a shortest action sequence that reaches a state violating a rule you wrote. Something like: Shell ✗ never-charged-twice [state] — pred returned false init {"status":"new","attempts":0,"hasDue":false} CREATE({}) -> {"status":"active","attempts":0,"hasDue":false} RENEW_CHARGE({"result":"5xx"}) -> {"status":"grace","attempts":1,"hasDue":true} RENEW_CHARGE({"result":"ok"}) -> {"status":"grace","attempts":2,"hasDue":true} That's a repro. You paste it into a test file, and you have a failing test in about ninety seconds. A real one: On a production SaaS subscription-billing machine, Polygraph flagged a disagreement on exactly one window: a 5xx from the payment processor during renewal moved the row to grace and marked it due, when the dunning path in the same codebase correctly treated 5xx as ambiguous. The next retry rotated the idempotency key. If the 503'd transfer had actually settled, the customer got charged twice. A human reviewer had found the same bug by hand; five independent model-derived readings of the source landed on it blind. And in the controlled seeded-bug eval, the split is worth knowing: replaying real traces against the derived spec found 0 of 5 seeded bugs. Model checking found 5 of 5, with counterexamples. Trace replay tells you whether to trust the model. Model checking is where the bugs actually are. What It Actually Does Three artifacts, all diffable, all in your repo: 1. contract.json — the scope. Which state fields matter, which actions the machine accepts, what data each action can carry, which states are terminal, and lang is the language your source is written in. 2. A spec — a JavaScript model of your code, written by an LLM from your source (whatever language that is). It's a strict SAM v2 module: every action it ignores has to say why via reject(reason), it can't hide bookkeeping state, and it declares its own action/data domains — so the checker knows what to explore with zero config. Several specs are generated independently and vote, so one bad generation doesn't decide anything. JavaScript export const stateInvariants = [ { name: 'locked-only-at-limit', pred: (s) => s.status !== 'locked' || s.attempts >= 3 }, ]; export const transitionInvariants = [ { name: 'expired-never-verifies', pred: (pre, action, data, post) => !(action === 'ATTEMPT' && data?.expired) || post.status !== 'verified' }, ]; 3. invariants.mjs — your rules, as plain JS predicates: This part is yours and can't be automated away. Code with a bug is a perfectly faithful description of the wrong behavior. Invariants are where your intent enters the system. Then two checks run. Replay asks "is the spec faithful?" Real traces ({pre, action, data, post} windows, captured by wrapping your dispatch once) are replayed against each spec, with positive and negative controls proving the harness can tell good from bad. Model check asks "where are the bugs?" It iterates the faithful spec exhaustively from init against your invariants and prints the shortest path to every violation. The Caveats "Exhaustive" means exhaustive over the finite (action, data) domain declared in your contract. A machine whose behavior depends on unbounded counters or arbitrary strings is checked only at the representative values someone chose. That's the standard TLA+ modeling move, and the gap between declared domain and real data is real.It's a consistency check, not a proof. A clean run means your code's observable behavior matches an independent reading of its own source. Nothing more.Every finding is a lead to investigate, not a verdict. There is no triage step that discharges "real invariant break with no observable consequence."It's experimental and not peer-reviewed. Don't make it your only safeguard on safety-critical code. API Key and Cost Only three things call the Anthropic API: spec generation, code authoring (polygen), and polynv's optional headless invariant harvest. You need ANTHROPIC_API_KEY in your environment for those, including inside Claude Code, where the skills shell out to the same scripts and do not use your session credentials. Ballpark, on a typical machine: you runkey?costverify.mjs --source … (generate + replay)yes~$0.50polygen.mjs --intent … (author new code — JS/TS output only)yes~$2replay saved specs, model check, --tla, polyvers, polynv, polyrunno$0 That second row is the load-bearing one. Everything that checks (replay, the exhaustive model check, version gating, the mutation grade, TLC escalation) is keyless, local, and deterministic on Node ≥ 20. Which is precisely what makes CI viable: you commit the spec, and the gate re-runs it on every merge request for free. No key in CI, no per-MR API bill, no nondeterminism in your pipeline. That gate is polygate, and there's a GitLab reference implementation at <POLYGATE_GITLAB_URL> — a .gitlab-ci.yml you can copy that runs corpus validation, replay, and the model check against your committed artifacts and fails the MR on a violation. (Contrast: Specula, the closest comparable agentic TLA+ pipeline, reports a median of $57 and 3.7 hours per system. Excellent tool, structurally can't run on every MR.) Getting Started Prerequisite, and it's a hard one: The stateful code has to be runnable in isolation, because traces are ground truth from the code actually executing. A clean step boundary: a dispatch, reducer, or handler, from experience, Claude will refactor it easily for you. If it only runs against a live DB or device, stand up doubles first (in Claude Code, the agent will build them). Note this is the only place your language matters, and only for convenience: the bundled withTracing / tapReducer helpers are JS, so a Go or Python machine means writing the {pre, action, data, post} NDJSON lines yourself. It's about ten lines. Zero-cost first (no key, five minutes): Shell git clone https://github.com/cognitive-fab/polygraph cd polygraph && npm test # validates the bundled corpus, runs the controls npm run verify:turnstile-v2 # replays bundled specs — see the output shape Then on your own machine, as a plugin: Shell /plugin marketplace add cognitive-fab/polygraph /plugin install polygraph@polygraph …and just ask: "verify this state machine", or /polygraph:polygraph for the guided end-to-end run (Claude drafts the contract, instruments the boundary, captures traces, runs controls, triages with you). Trace capture is historically what made this expensive; it's the step the agent now carries. Or plain CLI, no Claude Code: Wrap your dispatch once, projecting only the contract's observable keys (JS shown; in another language, emit the same NDJSON shape by hand): JavaScript import { withTracing } from '<plugin>/scripts/instrument/trace-emitter.mjs'; const dispatch = withTracing( rawDispatch, () => ({ status: m.status }),'traces/s1_normal.ndjson' ); Note --source takes your real file, in your real language: Shell node scripts/validate_corpus.mjs contract.json traces/ # no key node scripts/verify.mjs --contract contract.json --source src/machine.ts \ --traces traces/ --model opus-5 --n 5 --out out/ # key, ~$0.50 That writes out/findings.md and the generated specs to out/specs/. Commit the winning one, and from then on the loop is free: Shell node scripts/check.mjs --spec out/specs/spec_0.js --contract contract.json \ --invariants invariants.mjs # no key, forever There's no default model: pass --model. Use opus-5 or better; deriving a faithful transition function is a hard reasoning task and lighter models don't clear the bar. If you see empty specs, you lowered --max-tokens below what the reasoning block needs; put it back to 32000. Apache-2.0. The method is written up in arXiv:2607.05076. Your test suite is a sample. This is the census.

By Jean-Jacques Dubray
Incident Management and the Rise of AI SRE Agents
Incident Management and the Rise of AI SRE Agents

Over the past year, I've been rebuilding parts of an incident response stack for a client, and the biggest surprise wasn't the AI features themselves. It was how much of the underlying workflow had to change to make those features useful. You can't just bolt an LLM onto a 2015-era ticketing tool and call it AIOps. The queue structure, the alert taxonomy, even the way runbooks are written all need to change. I've written before about the agent side of this shift, in AI Agent Architectures: Patterns, Applications, and Implementation Guide and Observability and DevTool Platforms for AI Agents. This two-part series is the other side of that coin: what happens when you point those same agent patterns at your own production systems instead of at somebody else's AI application. Same reasoning loop, different target. This first part covers incident management specifically, including a category I skipped in my earlier tool roundups: dedicated "AI SRE" agents like Traversal, Resolve.ai, and Cleric, which behave differently from the AIOps platforms most of us grew up with. Part 2 goes past incidents into ITOps, chaos engineering, SLO management, on-call toil, and the rest of what fills an SRE's week. A note on the numbers below: vendor-reported accuracy and MTTR figures in this space move fast and come from the vendors themselves. I've flagged those clearly rather than presenting them as independently verified benchmarks. Where SRE Pain Actually Lives Before getting into tools, it helps to remember what SREs spend their time on. Most postmortems I read over the years have the same three complaints: Too many alerts, not enough signalCorrelating five different dashboards to find one root causeWriting the same postmortem summary for the fourth time this quarter None of these are new problems. What's new is that large language models are actually decent at the second and third ones, if you feed them clean data, and a newer crop of agents is starting to chip away at the first one too. The Traditional Incident Pipeline Here's roughly what an incident used to look like before AI got involved, at most mid-size shops I've worked with: Traditional incident pipeline Every arrow in that diagram is a human doing manual correlation work. That's fine when you have ten services. It falls apart at three hundred, and it's part of why I keep coming back to the point I made in Infrastructure as Code: How Automation Evolved to Power AI Workloads: scale problems in ops rarely get solved by hiring more people to stare at more dashboards. Where AI Fits Into the Pipeline Today The shift isn't "AI replaces the engineer." It's AI collapsing steps B through F into something closer to a single triage step, with the engineer reviewing a proposed root cause instead of hunting for one from scratch. AI-aided pipeline Notice the engineer never disappears from this diagram. They just move from being the one who does the correlation to the one who checks the correlation. That distinction matters, because it changes what you hire and train for. I made a version of this same argument about production-grade agents generally in the Shipping Production-Grade AI Agents refcard: an agent needs a human review layer, or you're just moving risk around instead of removing it. If you want a deeper look at how these review loops are actually structured under the hood, I broke that down recently in Loop Engineering: The Layer After Prompt, Context, and Harness Engineering. Incident Management: What Changed A few concrete things have improved in incident management tools over the last two years: Alert correlation got better. Tools like BigPanda, Moogsoft, and PagerDuty's AIOps features now cluster related alerts using pattern recognition instead of static rules. A database timeout, three downstream service errors, and a spike in 500s used to show up as four separate pages. Good correlation engines now group them as one incident with a suggested cause. Similar-incident retrieval works reasonably well. If your org has a decent history of past incidents with clean postmortems, tools can now surface "this looks like INC-4471 from March" with real accuracy. This only works if your postmortem data isn't garbage, which is a bigger blocker than people admit. Draft postmortems save real time. Not because the AI writes a good postmortem on the first try, but because staring at a blank page is the slowest part of writing one. A rough draft built from the incident timeline, Slack thread, and metrics gives engineers something to edit rather than create. What's newer, and worth its own section, is a class of tools that don't just correlate what you already collected. They go get new evidence during the incident, the way a senior engineer would. The New Category: Dedicated AI SRE Agents This is the part of the landscape that's moved fastest since I last wrote about agent tooling. A handful of startups have built agents whose entire job is investigating production incidents autonomously, not just clustering alerts that already exist. Traversal leans on causal machine learning rather than a general-purpose LLM wrapper. Instead of pattern-matching against similar past incidents, it builds a model of causal dependencies across your services and traces the actual chain of cause and effect, down to the specific deploy or config change that started the failure. It reports strong root-cause accuracy in production at large enterprises and is used for both alert triage and live incident investigation. The pitch is narrower than "full AIOps platform," and that narrowness is the point. Resolve.ai takes a broader angle. It was built by the team that created OpenTelemetry, and it positions itself as an agentic teammate across the whole production lifecycle: investigating incidents, but also touching capacity questions, config drift, and guided code changes. Where Traversal is a specialist in root cause, Resolve.ai is closer to a generalist you'd loop in on almost anything production-related, with the incident work as the anchor use case. Cleric sits in a similar space to both, with a specific focus on autonomous alert triage. It runs a multi-source investigation the moment an alert fires, pulling metrics, logs, traces, and deploy history in parallel, and returns an evidence-backed hypothesis before an on-call engineer has finished opening their second dashboard tab. It runs with read-only access by default, which matters a lot for teams still building trust in the category, and it was named a Gartner Cool Vendor in AI for SRE and Observability in 2025. That "read-only by default" design decision is exactly the kind of guardrail I argued for in Trust No Agent: How to Secure Autonomous Tools on Your Machine: an agent's blast radius should be a deliberate design choice, not an afterthought. Causely and NeuBird round out the space with slightly different angles: Causely focuses on causal reasoning to find the single root cause behind a storm of cascading alerts, and NeuBird targets enterprise IT environments with LLM-driven telemetry analysis at large scale. Here's roughly where an AI SRE agent sits in the pipeline compared to the AIOps correlation tools from the last section: AI SRE agent pipeline The key difference from the earlier diagram: this agent isn't just correlating signals you already collected in a dashboard. It's actively going out and querying your systems the way a human on-call engineer would, forming a hypothesis, testing it, and either confirming or discarding it before it ever pages a person. That's a meaningfully different capability than clustering alerts by similarity, and it's why this category gets its own row in any serious comparison. If you're weighing whether to build this kind of investigation loop yourself versus buying one of these platforms, it's worth reading MCP vs Skills vs Agents With Scripts: Which One Should You Pick? first, since the architecture decision behind "agent that calls tools live" versus "agent with a fixed skill set" applies just as much to SRE tooling as it does anywhere else. What This Actually Brings to the SRE Persona It's worth being specific about what changes for the person on-call, not just what the vendor deck claims: Fewer 2 a.m. investigations that start from zero. The agent has usually already ruled out the obvious suspects by the time a human looks at the page, so the engineer starts from a hypothesis instead of a blank terminal.Less tool-hopping. A lot of incident time isn't spent thinking; it's spent switching between Datadog, Grafana, the CI pipeline, and Slack. An agent that queries all of them in parallel removes a genuinely tedious chunk of the job.A written trail for free. Because the agent's investigation is itself a structured log of what it checked and why, you get a decent postmortem skeleton as a byproduct, not a separate task.A new failure mode to watch for. Engineers can start trusting the proposed root cause without checking the evidence trail, especially under pager pressure. That's a habit worth actively training against, not assuming away. None of this replaces the on-call engineer's judgment. It changes the shape of their shift from "gather evidence, then decide" to "review evidence, then decide," which is faster but only as trustworthy as the evidence the agent actually gathered. Comparing the Tool Landscape Here's how some of the major players stack up on where they've actually invested in AI, versus where it's mostly a checkbox feature. I've split this into two tables, because lumping AIOps correlation platforms in with dedicated AI SRE agents hides a real difference in what these tools do. Established AIOps and incident platforms: ToolAlert CorrelationRoot Cause SuggestionAuto-Drafted PostmortemsPredictive CapacityOwnership / StatusPagerDuty (AIOps)StrongModerateYesLimitedIndependent, public companyMoogsoftStrongStrongNoNoAcquired by Dell Technologies (2023)BigPandaStrongModerateLimitedNoIndependent, privateDatadog (Bits AI)ModerateStrongYesModerateBuilt in-house by DatadogServiceNow (Now Assist)ModerateModerateYesStrong (ITOps)Built in-house by ServiceNowDynatrace (Davis AI)StrongStrongLimitedStrongBuilt in-house by Dynatraceincident.ioModerateLimitedYesNoIndependent, privateRootlyModerateLimitedYesNoIndependent, private Dedicated AI SRE agents: ToolCore ApproachActs Autonomously?Best FitFounded / BackingTraversalCausal ML across dependency graphInvestigation autonomous, remediation guardedTeams with strong existing observability wanting sharper RCA2023, Sequoia and Kleiner PerkinsResolve.aiBroad agentic reasoning over code, infra, telemetryInvestigation autonomous, remediation opt-inTeams wanting one agent across incidents, capacity, and config2024, Greylock-led seedClericMulti-source parallel investigationRead-only by defaultTeams new to AI SRE agents, wary of write access2024, Zetta Venture PartnersCauselyCausal reasoning on cascading alertsInvestigation onlyEnvironments with alert storms and unclear blast radiusPrivate, early stageNeuBirdLLM-driven telemetry analysis at scaleInvestigation, guided remediationLarge enterprise IT environmentsPrivate, early stage A caveat worth stating plainly: I haven't run rigorous side-by-side benchmarks on all of these, and vendor claims move faster than reality, especially in the AI SRE agent table where most of these companies are one to three years old and evolving month to month. Treat both tables as a directional map, not a scorecard, and validate against your own alert volume before picking one. Deterministic AI vs. Generative AI in These Tools This distinction gets muddled in vendor marketing, so it's worth separating clearly. AspectDeterministic / ML-based (older AIOps)Generative AI (LLM-based, newer)ApproachStatistical pattern matching, clustering, anomaly detectionLanguage model reasoning over logs, tickets, chat history, and live queriesPredictabilityHigh, same input gives same outputLower, outputs can vary between runsStrengthCorrelation, anomaly detection at scaleSummarization, hypothesis generation, natural language explanation, draftingWeaknessPoor at explaining "why" in plain languageCan hallucinate a plausible-sounding but wrong root causeWhere it shows upDynatrace Davis AI, Moogsoft's original correlation engineDatadog Bits AI, ServiceNow Now Assist, Traversal, Resolve.ai, ClericTrust level neededCan often auto-remediateNeeds human review before action Most modern platforms now run both in tandem: the deterministic layer does the anomaly detection and correlation, and the generative layer explains it in plain English, forms hypotheses, and drafts the writeup. That combination is doing more real work than either piece alone, and it's basically the same pattern I described for agent observability generally in the AI agent architectures piece linked earlier: a fast, boring, reliable layer underneath a slower, flexible reasoning layer on top. Where We're Headed in Part 2 Incident response gets the spotlight because it's the loudest part of the job, but if you track where an SRE's actual week goes, a lot of it isn't firefighting at all. It's chaos testing, SLO math, on-call scheduling, and the slow grind of writing and maintaining runbooks nobody reads until 3 a.m. In Part 2, I'll walk through where AI is showing up in ITOps specifically, and then go further into chaos engineering, SLO and error budget management, on-call toil reduction, and capacity planning, the quieter parts of the job that determine whether the incident tools in this article even have a fighting chance.

By Vidyasagar (Sarath Chandra) Machupalli FBCS DZone Core CORE
I Got Tired of Copy-Pasting Microfrontend Boilerplate, So I Built a Bridge
I Got Tired of Copy-Pasting Microfrontend Boilerplate, So I Built a Bridge

When we started to work on microfrontend migration on one of our projects, the architecture looked great on paper (like always): one host shell, several remote apps, and teams could deploy independently on their own timelines. But in practice it wasn't so clean. One part kept getting on my nerves: actually mounting remote React components inside the host. Each microfrontend came with the same glue code. Load the remote bundle, create a React root, render the component, keep track of the mounted instance, push updated props into it when the host re-renders, and clean up listeners on unmount. And do not forget to handle load failures. It wasn't especially hard code. But it was just the kind of code nobody wants to repeat. Another problem is type safety, which had a habit of disappearing exactly where I wanted it most. Inside the remote, TypeScript understood the component props perfectly. But at the host boundary, that often collapsed into unknown and as any. If a remote added a required prop or renamed an existing one, the host usually did not find out from the compiler. After doing this a few times across different projects, I decided the pattern deserved a real abstraction instead of one more copy-pasted wrapper. What I Wanted It should be part of my toolkit package and shouldn't be really hard. Something much more practical. The goal was simple: Remove repetitive host-side boilerplateKeep prop types across the host/remote boundaryWork with separate bundles and separate React rootsAvoid shared stores, global registries, and code generationFit into an existing Module Federation setup without changing how remotes are versioned or deployed That idea transformed to @mf-toolkit/mf-bridge. The Base The package has two parts: one wrapper on the remote side, and one host component that takes care of the integration. On the remote side, you define the entry once: TypeScript import { createMFEntry } from '@mf-toolkit/mf-bridge/entry' import { CheckoutWidget } from './CheckoutWidget' export const register = createMFEntry(CheckoutWidget) On the host side, you render the bridge where the remote should appear: import { MFBridgeLazy } from '@mf-toolkit/mf-bridge' <MFBridgeLazy register={() => import('checkout/entry').then(m => m.register)} props={{ orderId, userId } fallback={<CheckoutSkeleton/>} /> That’s all. With MFBridgeLazy, the host doesn’t have to deal with all the hassle of loading things on demand, setting up the root, updating stuff, cleaning up, or handling event listeners — the tool does it all. Plus, because the register function has clear types, the host can automatically figure out what props the remote component needs. If the remote component suddenly needs a new prop, you’ll see a TypeScript error right away during development, not after the app is already live and causing problems. How Prop Updates Travel This was the part I wanted to keep as boring and predictable as possible. Once a remote component is mounted, it lives in its own React root. That means the host cannot simply re-render it as if it were a normal local child. The host still needs a way to send updated props into that remote tree every time its own state changes. There are plenty of ways to solve this: shared stores, shared context, global event buses, custom registries. I wanted the smallest possible mechanism that stayed local to each mounted microfrontend. So `mf-bridge` uses the one thing both sides already share: the mount element. When the host re-renders with new props, the bridge dispatches a `CustomEvent` on that specific DOM element. The remote listens to events on that same element and re-renders with the new props. That is it. I like this approach for a few reasons. First, it is naturally isolated. If you have several microfrontend slots on the same page, each one has its own mount element, so updates do not bleed across instances. Second, it does not need a shared module graph or global state container just to move props around. Third, it keeps the contract very explicit: the host owns the mount point, and the props, and the remote owns how it renders them. Internally, the package wraps this in a small typed DOM event bus, but consumers do not really need to think about those details. Why This Helped More Than Just Saving Lines of Code The obvious benefit is less boilerplate. If a page has five remote slots, I no longer end up with five slightly different wrappers all doing the same lifecycle work. But the bigger benefit is moving problems earlier in the process. Before this, the host/remote boundary was often exactly where type information got blurry. That made one of the most important contracts in the system feel surprisingly fragile. A remote could evolve, and the host would not always know it had fallen out of sync. With mf-bridge, prop inference flows from the remote entry to the host usage. That changes the feedback loop. A contract mismatch becomes a compile-time problem instead of an incident report. There is also a reliability benefit in the lifecycle handling. The package takes care of the repetitive, easy-to-forget parts: Lazy loading with a fallback UIClean mount and unmount behaviorProp streaming on re-rendersListener cleanupError handling when the remote fails to loadOptional preloading and retry behaviorOptional hooks for setup and teardown on the remote side when you need DI or per-mount initialization None of these features are individually groundbreaking. The value is that they come together in one small, reusable bridge instead of being re-implemented in every host wrapper. The Cases I Wanted to Be Sure About When the basic version started to work, I spent a bit more time on some of the scenarios that usually make microfrontend wrappers fragile. One of those cases was multiple instances of the same remote on a single page — a widget in the main content area, a compact version in a sidebar, or the same remote mounted in a few different places. I wanted to make sure what updates stayed local to the exact mount point instead of leaking. Using the DOM element itself as the transport turned out to be a very practical way to preserve that isolation. Another important case was failed loading. I didn't want the host to end up with a blank hole in the UI just because a remote bundle failed on the first attempt. That is why the bridge supports fallbacks, preloading, and retry behavior. I think that kind of thing makes an integration feel solid. And sure, we should not forget about what happens when the problem is rendering. If a remote drops during render, I do not want that failure to destabilize the whole host page. So error handling became part of the design too: we keep the failure contained to the mount point, surface the error to the host, and make recovery possible when new props arrive. Then there is setup and unmount — that case is covered, too. Where It Fits Compared to React.lazy or Portals This package is not a replacement for React.lazy, and it is not trying to be cleverer than React. If your component lives in the same bundle and the same React tree, React.lazy is still the natural tool. If you just want to render into a different DOM node inside the same tree, portals are great. mf-bridge is for the awkward case those tools do not cover well: a component living across a Module Federation boundary, loaded from a separate bundle, mounted into its own React root, but still expected to behave like a first-class part of the host page. That is the gap I wanted to close. A Small Package, Not a New Platform I also cared quite a bit about keeping the package lightweight. It has zero production dependencies and uses the browser's native CustomEvent API for prop streaming. In practice, that means less surface area, fewer moving parts, and one less utility layer to debug when something goes wrong. The goal was never to build a microfrontend platform. It was simply to remove a recurring nuisance and make the host/remote boundary feel safer. Sometimes that is enough to justify a package. I published it as @mf-toolkit/mf-bridge. Repository, docs, and examples: github.com/zvitaly7/mf-toolkit. If you are working with Module Federation and you already have a small pile of hand-written wrappers around remote React components, this may save you some time. And if you have solved the same problem in a completely different way, I would genuinely be curious to compare notes.

By Vitaly Zheltko
Building Internal Developer Platforms as Products: A Practical Guide for IDP Architects
Building Internal Developer Platforms as Products: A Practical Guide for IDP Architects

Why Most Platforms Fail to Become Products Many companies are heavily investing in internal developer platforms (IDPs) with the expectation that they will speed up delivery and governance, and increase developer productivity. Despite significant investment in Kubernetes, CI/CD, observability, security tooling, and cloud infrastructure, many platforms struggle to gain adoption. The reason is simple: they are built and operated like infrastructure projects, not products. Infrastructure teams are often very focused on technical excellence: automation, scalability, reliability, and compliance. Developers, on the other hand, are interested in a different goal — getting their applications into production quickly and safely without having to go through so much complexity. IDP is successful when developers choose it voluntarily because it makes their lives easier. That shift requires platform architects to think less like infrastructure engineers and more like product managers. Building an IDP is like operating an airport. Nobody travels because they love airports. They travel because they want to reach a destination efficiently. Similarly, developers do not care about Kubernetes clusters, pipelines, secrets management, or observability stacks. They care about shipping features to customers. The platform's job is to make the journey smooth, fast, and safe. This article explores the core practices that differentiate successful product-centric platforms from infrastructure-centric ones. Practice 1: Start With Developer Journeys, Not Technology Choices Imagine constructing a shopping mall by selecting elevators, security systems, and air-conditioning units before understanding customer traffic patterns. The result is often technically impressive but operationally frustrating. The same happens with developer platforms. Architects should first map the customer journey (developer journey) before designing platform capabilities. Many platform initiatives begin with questions like: Which Kubernetes distribution should we use?Which GitOps framework is best?Which CI/CD tool should be standardized? These are important questions, but they should not be the starting point. Successful platform architects begin by understanding developer workflows: How does a new service get created?How long does environment provisioning take?Where do deployment delays occur?What causes support tickets?Which activities are repetitive and manual? The goal is to identify friction and eliminate it. Organizations using platforms based on technologies like Red Hat OpenShift, IBM Cloud Kubernetes Service, or other cloud-native platforms have found that developers adopt only when the platform team focuses on reducing the friction in workflow rather than adding more infrastructure features to the platform. Practice 2: Create Golden Paths, Not Golden Handcuffs A highway encourages drivers to use the fastest route while still allowing exits when necessary. Successful IDPs behave like highways. Developers naturally choose the Golden Path because it is easier and safer than building everything from scratch. One of the most powerful concepts in modern platform engineering is the Golden Path. A Golden Path provides: Recommended architecturesStandard deployment patternsPre-approved security controlsBuilt-in observabilityAutomated CI/CD workflows Developers should be able to move fast along a paved road while retaining flexibility for unique requirements. Platform teams that leverage services from cloud provider environments often realize that standardized self-service templates drive significantly higher adoption than restrictive governance models. Practice 3: Make Self-Service the Primary Interface Every banking transaction once required a visit to a physical branch. Today, customers expect to do everything from a mobile app. Developers hope for the same experience from inside their own software. Nothing kills developer productivity faster than dependency queues. Consider a common case of dependency queues. Open a ticket for infrastructure.Wait for approval.Wait for provisioning.Request secrets.Request monitoring.Request deployment access. Weeks can pass before development even begins. Modern platforms must provide self-service experiences where developers can do the following without opening tickets. Create environmentsProvision databasesConfigure pipelinesAccess observability dashboardsRequest infrastructure resources An IDP should function like a digital banking application—secure, streamlined, and available on demand. Below is the Product-Centric IDP reference architecture. Developers consume platform capabilities through self-service experiences, while the platform embeds security, observability, governance, and delivery capabilities and exposes them through Golden Paths. Practice 4: Treat Platform APIs as Products A power drill might have sophisticated engineering in it. Users judge it by a very simple standard: “Can I drill a hole fast and reliably?" Many platform teams are focused on infrastructure automation and not developer experience. Each API, template, workflow, and portal interaction is a product interface. Questions worth asking include: Is the API predictable?Is documentation clear?Are error messages actionable?Is onboarding intuitive?Can developers discover capabilities easily? Developers evaluate IDPs the same way. They are not interested in the complexity underneath. They care about usability. This principle is especially important when integrating observability services, cloud provisioning layers, or deployment automation platforms. For example, IBM Cloud's managed services can significantly simplify operational complexity, but value is realized only when developers experience that simplicity through intuitive platform workflows. Practice 5: Build Observability into the Platform, Not Around It Imagine when you are driving a car without any speedometer, fuel gauge or warning indicators. You may still reach your destination but the risk increases dramatically. Observability is the dashboard for software systems. Observability is often treated as an afterthought. A team deploys an application and later attempts to add: MetricsLogsTracesDashboardsAlerting This approach creates inconsistency and operational blind spots. Platform teams should embed observability from day one. Every service created through the platform should automatically include: Logging standardsDistributed tracingMetrics collectionHealth monitoringService dashboards Whether organizations use IBM Cloud Observability, Instana, OpenTelemetry, Prometheus, Grafana, or other solutions, the platform should make observability automatic rather than optional. Practice 6: Make Security Invisible but Ubiquitous When entering a modern office building, people rarely think about security. Access badges, surveillance, and emergency controls are built into the environment — the building is secure without requiring employees to become security experts. The same principle applies to IDPs. In immature environments, security is seen as a series of checkpoints, review meetings, manual compliance approvals, vulnerability assessments, and audit evidence collection. Developers find it as friction because it arrives late in the delivery lifecycle. Traditional security models operate as gates. Platform-centric security operates as guardrails. The objective is not fewer security controls — it is fewer manual interactions. Build Secure-by-Default Golden Paths Every new service created through the platform should automatically inherit: Secure CI/CD pipelines with dependency and container image scanningSecret detection and policy enforcementAccess control standards and audit loggingEncryption best practices Automate Policy Enforcement Manual compliance verification is one of the biggest sources of deployment delays. Platform teams should adopt policy-as-code (PaC) approaches that automatically validate deployment configurations, infrastructure standards, and regulatory controls. Instead of asking, "Did someone review this configuration?" the platform asks, "Does this configuration satisfy our policies?" Reduce Security Cognitive Load Developers should not need deep expertise in every security domain. The platform should abstract identity management, secrets management, certificate management, and vulnerability remediation workflows—particularly in hybrid and multi-cloud environments where security complexity grows rapidly. A useful measure of progress: the percentage of security controls inherited from the platform versus manually implemented by application teams. The higher the inheritance rate, the lower the cognitive load. Practice 7: Measure Platform Success Like a Product A gym owner does not measure success by counting treadmills—they measure it by member outcomes. Platform teams should apply the same logic. Traditional infrastructure metrics like cluster utilization, pipeline counts, and resource consumption tell you whether the platform is running. They do not tell you whether it is working for developers. Product-oriented platform teams focus on: Developer satisfactionPlatform adoptionTime to first deploymentDeployment frequencyLead time for changes If developers still circumvent the platform, no amount of technical sophistication matters. The Developer Experience Scorecard Measuring developer experience requires balancing sentiment, effort, and adoption. High-performing platform teams track four key measures: Metric What It Measures How to Collect Developer Satisfaction Score (DSS) Overall platform sentiment Quarterly survey, 1–10 scale Platform NPS Willingness to recommend the platform "How likely are you to recommend this platform?" scored 0–10 Ease-of-Use Score How intuitive common workflows feel Per-task rating, 1–5 scale Developer Effort Score How much work is required to achieve an outcome Survey question on effort per task Together, these reveal not just whether developers are using the platform but whether they genuinely value it. Satisfaction Is a Leading Indicator Most delivery metrics lag behind—deployment frequency (e.g., lead time, incident count) and other metrics. Developer satisfaction is a leading indicator. Developers discover friction long before it is observable from the data. A declining DSS today will result in a decline in productivity and adoption tomorrow. Listening early allows platform teams to respond before problems grow into organizational challenges. The real measure of success is not how many developers use the platform—it is how they feel while using it. The IDP Health Dashboard High-performing platform teams monitor a balanced set of metrics across four categories: Category Metrics Sentiment DSS, Platform NPS, Ease-of-Use ratings Adoption Golden Path adoption, self-service usage, onboarding rates Friction Support ticket volume, documentation search failures, manual approval requests Productivity Time to First Deployment (TTFD), environment provisioning time, lead time for changes A platform succeeds not when developers are forced to use it, but when they prefer to use it. Practice 8: Reduce Cognitive Load Relentlessly The automotive industry spent decades simplifying the driving experience so drivers could focus on reaching their destination rather than understanding the mechanics of their vehicles. IDPs should do the same. As organizations evolve into cloud-native architectures, developers are expected to navigate containers, Kubernetes, CI/CD, IaC, security policies, service meshes, observability tools, and compliance requirements all at once. Each one solves a very important problem individually. As a whole, they overwhelm developers and take focus away from developing business capabilities. A successful platform is not one that exposes every infrastructure capability. It is one that hides unnecessary complexity while providing simple, intuitive paths to outcomes. The goal of platform engineering is not to eliminate complexity. It is to absorb complexity so developers don't have to. Common indicators of excessive cognitive load: Developers struggling to find documentationFrequent support requests for routine tasksLong onboarding times for new servicesMultiple handoffs between teamsTool sprawl across the engineering ecosystem Reduce Tool Sprawl Every tool a developer must learn introduces new interfaces, terminology, documentation, and configuration models. Platform teams should create a unified experience through a developer portal, service catalog, or platform API, that minimizes the number of decisions and interfaces developers encounter. Minimize Context Switching Every transition between tools, teams, or approval processes introduces cognitive overhead. Platform teams should ask: Can this be automated? Can these steps be consolidated? Can approvals be replaced with automated guardrails? The goal is fewer interruptions between code creation and deployment. Platform Teams Are Complexity Brokers Complexity never disappears — it moves. Organizations can either push complexity onto every development team, or centralize and manage it within the platform. High-performing platform teams choose the latter, absorbing operational, security, infrastructure, and compliance complexity so application teams can focus on features. Practice 9: Obsess Over Time to First Deployment The first experience developers have with a platform often determines whether they embrace it or avoid it. Imagine a shopping mall where opening a new store requires twelve forms, multiple approval queues, and manual setup of every utility. Store owners would go elsewhere. The best malls provide ready-made spaces where businesses can start operating almost immediately. Developer platforms should do the same. High-performing platform teams focus relentlessly on Time to First Deployment (TTFD) — the time between creating a service and successfully deploying it. The Biggest Contributors to Poor TTFD Bottleneck Root Cause Fix Manual infrastructure provisioning Ticket-driven approval chains Self-service IaC, service catalogs, platform portals CI/CD pipelines built from scratch No standard templates Pre-built, reusable pipeline templates Security reviews at the end Late-stage compliance gates Shift left — embed scans and policy checks in Golden Paths Observability setup delays Manual metrics/dashboard configuration Auto-provision logging, tracing, and health checks by default Too many decisions Choice overload at onboarding Provide Golden Paths with sensible defaults Measure Every Stage Stage Target Service creation < 5 mins Repository creation Automated Pipeline creation Automated Infrastructure provisioning < 10 mins First build < 5 mins First deployment < 15 mins Observability enablement Automatic TTFD = Provisioning Time + Setup Time + Approval Time + Deployment Time Many organizations discover that approval time is larger than all technical activities combined. The fastest platforms replace approvals with automated guardrails. Practice 10: Build a Platform Community, Not Just a Platform Team Cities flourish when residents contribute feedback and shape growth. Cities planned entirely from a central authority often struggle to meet citizen needs. IDPs are no different. The best platforms evolve through continuous collaboration. Platform teams should create feedback loops through office hours, community forums, developer councils, internal documentation reviews, and experience surveys. Developers become co-creators rather than consumers. Community Health Metrics Running community mechanisms is not enough — each one needs a way to know whether it is working. Track these six indicators to measure community health: Metric What It Measures Healthy Signal Monthly Active Community Members Developers engaging in forums, channels, or office hours Steady growth quarter over quarter Developer-to-Developer Answer Rate % of forum questions answered by non-platform-team members Above 40% indicates a self-sustaining community External Contributions per Quarter Pull requests or documentation edits from application teams Increasing trend Roadmap Items from Community Input % of platform backlog items originating from developer feedback Above 50% signals product-centric culture Office Hours Repeat Attendance Rate % of attendees who return across multiple sessions Above 60% indicates ongoing value Support Ticket Deflection Rate % of issues resolved via community before a ticket is opened Rising deflection reduces platform team toil The ultimate sign of a mature platform community is a change in how developers talk about the platform—from something that happens to them to something they help shape. Practice 11: Think in Products, Roadmaps, and Customer Value Smartphones succeeded because manufacturers continuously improved user experience. Customers did not buy phones because of processor specifications. They bought outcomes—better communication, productivity, and convenience. Developers adopt platforms for the same reason. The strongest indicator that a platform is becoming a product is a change in language. Instead of asking: What infrastructure should we standardize? Platform teams begin asking: What developer problems should we solve next? Which user journeys create the most friction?Which capabilities deliver the highest value?What does our product roadmap look like? Features matter only when they improve the developer experience. Practice 12: Design for Platform Reliability, not Just Application Reliability Imagine a city that invests heavily in building roads, bridges, and public transport for its citizens, but has no maintenance crew, no traffic monitoring, and no plan for when a bridge closes. The infrastructure exists, but without reliability commitments, citizens cannot depend on it. Internal developer platforms face exactly the same risk. Most platform engineering conversations focus on the reliability of applications running on the platform — uptime, error rates, latency SLOs for customer-facing services. What is rarely discussed is the reliability of the platform itself. Yet the platform is load-bearing infrastructure for every engineering team in the organisation. When the CI/CD pipeline degrades, every team's delivery stops. When the service catalog is unavailable, no new services can be provisioned. The platform's reliability is a multiplier — a single failure can simultaneously impact dozens of teams. Define Platform SLOs Before Developers Define Them for You Platform teams that do not define their own Service Level Objectives will find that developers define them informally — through frustration, workarounds, and loss of trust. Effective platform SLOs cover the experiences developers depend on most: Pipeline availability — what percentage of CI/CD pipeline executions succeed without infrastructure-related failures?Provisioning latency — how long does environment or resource provisioning take at the 95th percentile?Portal availability — is the developer portal and service catalog accessible during working hours?Golden Path build time — how long does a standard pipeline template take to complete? These are the experience metrics developers encounter every day. A platform team that publishes and tracks these SLOs operates as a reliable internal service provider. A team that does not is invisible until something breaks. IDP Maturity Model Stage Characteristics Infrastructure Platform Standardized infrastructure, clusters, CI/CD tooling Self-Service Platform Service catalogs, automation, infrastructure on demand Developer Platform Golden Paths, integrated observability and security, DevEx focus Platform Product Platform roadmaps, adoption metrics, developer satisfaction measurement Adaptive Platform Continuous feedback loops, AI-assisted operations, continuous platform evolution Most organizations do not start with a Platform Product. They evolve toward it. The goal of the maturity model is not to reach the highest stage overnight, but to identify the next set of capabilities that will improve developer experience and platform adoption. High-performing platform teams treat platform maturity as a journey rather than a destination. Assessing Your Current Stage To identify where your platform currently sits, ask three diagnostic questions: How do developers access platform capabilities today? If the answer is "by opening a ticket," the platform is at the infrastructure stage. If developers provision resources on demand without human approval, they are at the self-service stage or beyond.Do developers choose the platform voluntarily or use it because they must? Voluntary adoption driven by speed and simplicity signals a developer platform or platform product. Mandatory usage with frequent workarounds signals an earlier stage.Does the platform team maintain a product roadmap prioritized by developer feedback? A yes here is the clearest indicator of a platform product. The absence of a roadmap almost always reflects an infrastructure or self-service mindset. Moving to the Next Stage Each stage has a single dominant unlock that drives progression: Infrastructure → Self-Service: Replace ticket-driven provisioning with self-service automation and a service catalog.Self-Service → Developer Platform: Introduce Golden Paths that embed security, observability, and CI/CD by default.Developer Platform → Platform Product: Establish a formal platform roadmap, measure developer satisfaction (DSS, NPS), and treat developer feedback as a product backlog.Platform Product → Adaptive Platform: Build continuous feedback loops, introduce AI-assisted operations, and invest in platform telemetry that proactively surfaces friction before developers report it. The most common mistake is attempting to skip stages. Teams that build Golden Paths before self-service exists create well-designed paths nobody can access independently. Teams that adopt satisfaction metrics before Golden Paths exist measure friction without the tools to address it. Progress through the stages in order. The IDP Architect's Checklist Before launching any new platform capability, ask: ✅ Does this feature remove friction from a developer workflow? ✅ Can developers access it through self-service? ✅ Is it aligned with a Golden Path? ✅ Is observability included by default? ✅ Is security built into the platform? ✅ Is governance automated rather than manual? ✅ Can success be measured through developer outcomes? ✅ Does it reduce cognitive load? ✅ Does it improve Time to First Deployment? ✅ Would developers choose this platform if they had alternatives? If the answer to several of these questions is "no," the capability is probably infrastructure-focused rather than product-focused. Final Thoughts The future of platform engineering is not about building more infrastructure. It is about delivering better developer experiences. The most successful IDPs combine the discipline of site reliability engineering (SRE), the automation of cloud-native technologies, and the mindset of product management. Whether your foundation runs on IBM Cloud, OpenShift, hyperscaler cloud services, or a hybrid environment, the winning formula remains the same: Treat developers as customers. Treat the platform as a product. Treat developer productivity as the ultimate business metric. When platform architects embrace this mindset, platforms stop being collections of tools and start becoming accelerators of innovation—and that's when platforms truly become products.

By Josephine Eskaline Joyce, Ph.D DZone Core CORE
AI in SRE: A Practical Autonomy Model for Self-Healing Infrastructure
AI in SRE: A Practical Autonomy Model for Self-Healing Infrastructure

Most SRE teams do not need another dashboard. They need a safer way to move from "something is wrong" to "we know what to do next." A model that detects anomalies is useful. A model that can touch production can also make a bad incident worse. That is where most conversations about AI in SRE become too optimistic for my taste. The hard part is not only detection. It is deciding how much autonomy the system should have, under which conditions, and with what blast-radius controls. I learned this while working on large-scale cloud services where one customer-facing symptom could turn into a flood of alerts. A degraded dependency might show up as latency in one service, retries in another, queue growth somewhere else, and CPU pressure downstream. During an on-call shift, that can look like five separate problems. Usually, it is one problem echoing through the stack. That experience changed how I think about self-healing infrastructure. The goal is not to build a system that blindly fixes everything. The goal is to build an operational control loop that can separate routine, low-risk recovery from incidents that still need human judgment. The model that has worked best for me is graduated autonomy: Let the system act automatically only when the action is well understood, reversible, and narrow in blast radius. For everything else, the system should collect evidence, recommend the next step, and keep humans in control. Why Static Alerts Stop Scaling Static alerts are not the enemy. I still want to know when disk usage is dangerous, error rates spike, or latency crosses a service-level threshold. But thresholds do not understand context. A CPU spike during a scheduled batch job may be normal. The same spike during steady-state traffic may be a retry storm. A latency increase in one region may be harmless during a controlled deployment, but suspicious if it appears across multiple availability zones with no recent change event. At small scale, engineers can carry that context in their heads. At enterprise scale, they cannot. Services emit hundreds of metrics across regions, dependencies, deployments, and customer paths. Eventually the team is no longer tuning alerts. It is negotiating with noise. In one rollout I was involved with, the most useful improvement was not adding more alerts. It was grouping alerts around dependency context and suppressing repeated downstream symptoms. The on-call experience became calmer because engineers could focus on the likely failure path instead of chasing every red graph independently. That is the kind of problem AI can help with. Not by replacing SRE judgment, but by organizing noisy signals into a more useful operational story. Detection Is Only the First Layer ML-based anomaly detection helps because it learns a service's normal operating shape instead of relying only on fixed thresholds. For cloud metrics, that usually means learning seasonality, traffic cycles, deployment windows, regional differences, and service-specific behavior. An LSTM autoencoder, isolation forest, or well-tuned statistical baseline can all be useful. I care less about the model family than the quality of the telemetry around it. A simple model trained on clean, consistent data will usually beat a sophisticated model trained on messy metrics. A practical anomaly pipeline usually looks like this: Collect metrics, logs, traces, and change events.Normalize them by service, region, dependency, and time window.Score each signal against its learned baseline.Group anomalies by dependency graph and recent changes.Produce an evidence bundle for automation or human review. Here is a simplified version of the scoring stage: Python from dataclasses import dataclass from typing import List @dataclass class MetricWindow: service: str region: str signal: str values: List[float] recent_deploy: bool = False @dataclass class AnomalyScore: service: str region: str signal: str score: float reason: str class BaselineModel: def expected_range(self, service: str, region: str, signal: str): # In production, this may come from a trained model, # feature store, or rolling baseline per service and region. return (0.0, 1.0) def score_window(window: MetricWindow, baseline: BaselineModel) -> AnomalyScore: low, high = baseline.expected_range( window.service, window.region, window.signal, ) latest = window.values[-1] if latest > high: distance = (latest - high) / max(high, 0.001) reason = f"{window.signal} above learned baseline" elif latest < low: distance = (low - latest) / max(abs(low), 0.001) reason = f"{window.signal} below learned baseline" else: distance = 0.0 reason = "within learned baseline" if window.recent_deploy and distance > 0: reason += " during recent deployment window" return AnomalyScore( service=window.service, region=window.region, signal=window.signal, score=min(distance, 1.0), reason=reason, ) The production value is not just the score. It is the metadata around it: ownership, dependency path, recent deploys, feature flag changes, customer impact, and whether the same pattern has appeared before. A single anomalous metric should rarely trigger remediation. Sustained anomalies across correlated signals are more trustworthy than one spike in one chart. Correlation Turns Noise Into an Incident Story During an incident, the useful question is not "Which graph is red?" It is "What changed first, and what depends on it?" That is where dependency-aware correlation becomes more useful than raw anomaly detection. A database issue may surface as API latency, retries, queue saturation, and CPU pressure. Without a dependency graph, every downstream service looks guilty. With one, the system can rank likely causes instead of handing the engineer a wall of symptoms. A useful correlation engine should look at topology, timing, change context, and customer impact. Which dependency failed first? Was there a deployment or config change? Which service is closest to the customer-facing error? The evidence bundle should be readable by a human. If the model says "root cause confidence: 0.86," that is not enough. It should also explain why. JSON { "candidate_root_cause": "identity-token-cache", "region": "example-region-1", "confidence": 0.86, "customer_impact": "elevated authentication latency for a subset of requests", "supporting_signals": [ "p99 latency above learned baseline for multiple consecutive windows", "cache hit rate dropped below its recent operating range", "downstream services showed retry growth after the initial cache anomaly", "no database saturation was observed", "no deployment was detected in the immediate incident window" ], "recommended_action": "drain_and_restart_one_cache_node", "estimated_blast_radius": "single node in a redundant pool", "rollback_plan": "keep node out of rotation if health checks fail after restart" } This is more useful than another alert. It gives the on-call engineer a starting hypothesis and the reasoning behind it. The Graduated Autonomy Model The most important design decision in self-healing infrastructure is not which ML algorithm to use. It is which actions the system is allowed to take. I divide remediation into three tiers. Tier 1: Fully Automated, Low-Risk Actions Tier 1 actions are safe, reversible, and narrow in blast radius. These are actions the system can execute without waiting for a human when confidence is high. Examples include restarting one unhealthy instance, scaling out a stateless service, draining one bad node, flushing a bounded cache, or shifting a small amount of traffic away from a degraded zone. The key phrase is bounded blast radius. Auto-remediation should not restart half the fleet, fail over a primary database, or disable a feature globally just because a model is confident. Confidence is not a substitute for safety. Before I put an action in Tier 1, I expect it to pass these checks: it is reversible, affected capacity is small, redundancy is healthy, there is no active global incident, the same action has not failed recently, rollback is defined, and health checks can verify success quickly. The first Tier 1 actions should be boring. Restarting one unhealthy node is not exciting, but it is exactly the kind of action that can be automated safely when the system has enough evidence. Tier 2: Automated Recommendation With Human Approval Tier 2 is where many real incidents live. The system may know what should happen, but the action still needs human approval. Examples include rolling back a deployment, disabling a feature flag, failing over a database, increasing capacity beyond a normal band, or changing regional routing. For Tier 2, the system should prepare the action, show the evidence, and ask for approval. The human should decide whether the action makes sense, not build the command during the incident. One pattern I have seen repeatedly: the slowest part of remediation is not always finding a likely cause. It is gathering enough confidence to take a risky action. When the system attaches deploy timing, error movement, affected endpoints, config changes, and rollback commands into one review card, the decision becomes easier. Tier 3: Human-Led With AI Context Tier 3 incidents are novel, high-risk, or ambiguous. The system should not execute remediation. It should help humans reason. This includes possible data corruption, multi-region cascading failures, security-sensitive incidents, conflicting signals across dependencies, low-confidence root-cause analysis, or any action with unclear rollback behavior. In Tier 3, the system's job is to summarize what it knows, what changed recently, which hypotheses are most likely, and which dashboards or runbooks are relevant. That alone can save time, but it keeps production control where it belongs. Architecture: A Control Loop, Not a Magic Button A practical self-healing system looks like a control loop with guardrails. Architecture diagram: Graduated autonomy model for self-healing infrastructure The important part of this diagram is the policy gate. Detection and correlation produce a recommendation, but the policy gate decides autonomy. Without that layer, "self-healing" becomes a risky automation script with an ML label attached. The policy gate should evaluate confidence, risk, blast radius, recent action history, service criticality, and rollback readiness. I would express that as policy-driven code: JSON from dataclasses import dataclass from enum import Enum from typing import List class Decision(str, Enum): AUTO_EXECUTE = "auto_execute" REQUEST_APPROVAL = "request_approval" HUMAN_LED = "human_led" @dataclass class RemediationProposal: action: str confidence: float blast_radius_percent: float reversible: bool rollback_defined: bool service_tier: str evidence: List[str] @dataclass class RuntimeContext: active_global_incident: bool recent_failed_action: bool healthy_redundancy: bool minutes_since_last_same_action: int TIER_1_ACTIONS = { "restart_single_instance", "scale_stateless_service", "drain_single_node", "flush_bounded_cache" } TIER_2_ACTIONS = { "rollback_deployment", "disable_feature_flag", "database_failover", "regional_traffic_shift" } def decide_autonomy( proposal: RemediationProposal, context: RuntimeContext ) -> Decision: if context.active_global_incident: return Decision.HUMAN_LED if context.recent_failed_action: return Decision.HUMAN_LED if not proposal.rollback_defined: return Decision.HUMAN_LED if proposal.action in TIER_1_ACTIONS: safe_enough = all([ proposal.confidence >= 0.90, proposal.blast_radius_percent <= 5.0, proposal.reversible, context.healthy_redundancy, context.minutes_since_last_same_action >= 30, len(proposal.evidence) >= 3, ]) return Decision.AUTO_EXECUTE if safe_enough else Decision.REQUEST_APPROVAL if proposal.action in TIER_2_ACTIONS and proposal.confidence >= 0.75: return Decision.REQUEST_APPROVAL return Decision.HUMAN_LED This is not drop-in production code, but the structure is the point: actions are classified, confidence is not the only input, and safety can override the model. In reliable systems, the model proposes; policy disposes. What I Measure Before Expanding Autonomy I would not start by asking, "Can we automate remediation?" I would start by asking whether the system's recommendations are trustworthy. Before allowing Tier 1 execution, I would track root-cause precision, false positives by service, recommendation acceptance, time to useful diagnosis, remediation success, rollback frequency, and any secondary incidents caused by remediation. The last two matter the most to me. A self-healing system that fixes one issue but creates another is not healing. It is moving the incident. My preference is to run in shadow mode first. Let the system detect, correlate, and recommend, but do not let it execute. Compare its recommendations against what engineers actually did. Once the system repeatedly recommends the same low-risk actions humans already take, graduate those actions into Tier 1. That is how trust gets built: not through a big launch, but through repeated correctness in narrow, well-understood situations. Lessons Learned From Building Toward Self-Healing The most useful lessons are not about model architecture. Clean telemetry beats clever models. If service names are inconsistent, regions are missing, logs are unstructured, and ownership metadata is stale, the model will struggle. Before debating LSTMs versus transformers, fix the telemetry pipeline. Change events are first-class signals. Deployments, config pushes, schema changes, and feature flag flips explain many anomalies. If the model cannot see change events, it will treat every incident like a mystery. Alert suppression is not the same as diagnosis. Reducing noise is useful, but the system must preserve the causal path. Suppressing duplicate downstream alerts only helps if the upstream root cause remains visible. Automation needs a memory. Every remediation should leave an audit trail: what was detected, what action was taken, what happened afterward, whether rollback was needed, and whether humans agreed with the recommendation. Start with boring actions. Restarting one bad instance is not glamorous. Draining one node is not a research breakthrough. But these are exactly the kinds of actions that make sense for early autonomy because they are repeatable, reversible, and easy to verify. Where LLMs Fit Large language models are useful in SRE, but I would not put them directly in the execution path for remediation. Their best role is communication and context assembly. An LLM can draft an incident summary, explain the evidence bundle, turn raw telemetry into a timeline, identify runbooks, and prepare a post-incident report. That saves time without giving the model direct control over production. The safer pattern is separation of responsibilities: ML or statistical models detect anomalies, graph correlation ranks likely causes, policy gates decide autonomy, deterministic automation executes approved actions, and LLMs summarize what happened. That separation keeps the high-risk parts deterministic and auditable while still using AI where it helps most. Final Thought Self-healing infrastructure is not about removing SREs from production. It is about removing the repetitive, low-risk work that slows them down during incidents. The best version of AI in SRE is not a magic system that fixes everything. It is a careful control loop: detect early, correlate intelligently, act only within policy, and learn from every outcome. If you are building toward self-healing, do not start with full autonomy. Start with evidence. Then recommendations. Then approval-based actions. Then, only after the system has earned trust, allow narrow automated remediation. That path is slower than the hype cycle, but it is much closer to how reliable infrastructure actually gets built.

By Shraddhaben Gajjar
Lift-and-Shift vs. Modernize: A Decision Framework for Enterprise Workloads
Lift-and-Shift vs. Modernize: A Decision Framework for Enterprise Workloads

One of the most consequential decisions in any enterprise cloud migration is deceptively simple to state and surprisingly hard to answer: do we move the workload as-is, or do we modernize it first? Having worked through cloud migrations across dozens of enterprise customers spanning both AWS and Azure. I can tell you this question rarely has a universal answer. The right path depends on the workload, the business context, and the maturity of the team inheriting it in the cloud. What follows is the decision framework I use when guiding customers through this choice. Understanding the Two Paths Lift-and-shift (also called rehost) means moving a workload to the cloud with minimal or no code changes. You are essentially taking an on-premises virtual machine (VM), an application server, or a database and running it on cloud infrastructure instead. Tools like Azure Migrate and AWS Migration Hub (Application Migration Service, or MGN) are purpose-built for this. Modernization is a broader term that can mean refactoring an application to use cloud-native services (databases-as-a-service, managed Kubernetes, serverless functions), re-platforming to a container-based architecture, or rebuilding from scratch as a microservices application. The spectrum between these two poles includes re-platforming, for example, moving a SQL Server workload to Azure SQL Managed Instance, which preserves the database engine behavior while offloading infrastructure management. This middle path is often underrated. The Core Tension Lift-and-shift is fast and low-risk. You can move a workload in weeks, not months. Your teams do not need to rearchitect anything. Applications continue to behave exactly as they did on-premises. The downside is that you carry your technical debt into the cloud. A poorly designed, resource-hungry application that cost you money on-premises will likely cost you more in the cloud, where idle compute is billed by the hour. You also miss out on cloud-native capabilities: autoscaling, managed resilience, and pay-per-use economics. Modernization promises better long-term economics and agility. But it is expensive up front, requires skill sets your team may not yet have, and introduces real delivery risk. Projects that start as modernization efforts frequently run over time and budget. The goal of a decision framework is to apply the right approach to the right workload, not to pick a single philosophy and apply it everywhere. Five Questions That Drive the Decision 1. What Is the Business Criticality of This Workload? Tier 1: Applications that directly generate revenue or are customer-facing warrant investment in modernization, especially if they have growth potential. The engineering effort pays back through scalability, resilience, and feature velocity. Tier 3: Internal tools, reporting systems, or legacy applications used by a handful of employees are strong lift-and-shift candidates. The cost of modernizing rarely justifies the benefit. A fast triage: Ask the application owner what happens if the application is down for four hours during business hours. The answer tells you a lot about where to invest. 2. Is the Application End-of-Life or Actively Developed? If an application is on a deprecation path, to be replaced in 18 to 36 months, lift-and-shift is almost always the correct call. You want the application in the cloud for consolidation, cost, or data center exit reasons, but you do not want to invest engineering resources in something you are going to retire. Conversely, if an application is actively developed and your engineering team ships features to it regularly, modernization has a compounding return. Every sprint benefits from cloud-native capabilities. 3. What Are the Licensing and Dependency Constraints? Some applications are locked to specific operating system versions, middleware versions, or third-party components that are not certified on modern platforms. A manufacturing execution system or a financial ledger application from 2008 may have an ISV (Independent Software Vendor) support contract that explicitly requires Windows Server 2012 R2. In those cases, your choice is not lift-and-shift versus modernization. It is lift-and-shift or do nothing. Azure and AWS both offer extended security update programs for legacy OS versions, making rehost viable even for older stacks. 4. What Are the Team's Skills and Capacity? Modernization is an engineering-intensive activity. If your team is composed primarily of infrastructure engineers skilled at VM management but with limited experience in Kubernetes, Terraform, or cloud-native PaaS (Platform as a Service) services, a forced modernization will stall. Honest capacity and skills assessment matters. I have seen organizations attempt to modernize a monolithic Java application to microservices while simultaneously running a datacenter migration. Both programs suffered. A phased approach often works better: lift-and-shift first to get out of the datacenter, then modernize workloads incrementally once the team is stable on the cloud platform. 5. What Are the Unit Economics Over a Three-Year Horizon? Run the numbers. This is non-negotiable. Tools like Azure's Total Cost of Ownership (TCO) calculator or AWS Pricing Calculator can model lift-and-shift costs quickly. For modernization, you will need to factor in engineering labor costs, which are often 3x to 5x the infrastructure savings in the first year. The business case shifts in favor of modernization when: The workload has high and variable traffic (autoscaling delivers real savings)The team plans significant feature development (cloud-native accelerates delivery)The current architecture requires expensive licensed middleware that PaaS services can replace The business case favors lift-and-shift when: The workload has predictable, flat traffic (reserved instances close the cost gap)Engineering capacity is constrainedThe migration is driven by a hard datacenter exit deadline A Decision Matrix Factor Favor Lift-and-Shift Favor Modernize Business criticality Low to medium High, customer-facing Development activity Stable / end-of-life Active development Technical debt Manageable High and growing Team skill set Infrastructure-focused App dev / cloud-native capable Timeline Hard deadline Flexible Licensing constraints ISV-locked Open or replaceable Traffic pattern Flat, predictable Variable, spiky The Re-Platform Middle Path Before forcing a binary choice, evaluate re-platforming for database workloads. Moving from SQL Server on a VM to Azure SQL Managed Instance, or from Oracle to Amazon RDS, is a lift-and-shift at the application layer and a modernization at the data layer. You eliminate OS patching, get automated backups, built-in high availability, and elastic scaling without refactoring a single line of application code in most cases. This is often the highest-return migration move available to enterprise customers and is underutilized because teams think in binary terms. What I See Go Wrong The most common failure mode is scope creep driven by modernization enthusiasm. A team scopes a lift-and-shift, then someone says “while we’re at it, let’s containerize it.” Twelve months later, the application is still not in production. The second most common failure mode is lift-and-shift without right-sizing. Teams migrate on-premises VMs 1:1 to cloud VMs without analyzing actual CPU and memory utilization. Azure Migrate’s performance-based assessments and AWS Compute Optimizer exist for exactly this reason. A VM provisioned at 16 cores on-premises is often running at 8% CPU utilization. Moving it as-is is leaving money on the table. Both mistakes are avoidable with a disciplined assessment phase before migration execution begins. Putting the Framework Into Practice In a typical enterprise migration engagement, I recommend the following sequencing: Discover and classify: Run an agentless discovery (Azure Migrate or AWS MGN) to inventory all workloads. Classify each by tier, development activity, and licensing constraints.Apply the decision matrix: Score each workload and assign a migration strategy: rehost, re-platform, or modernize.Sequence by risk: Start migrations with lower-criticality, lower-complexity workloads to build team confidence on the target platform.Right-size before you migrate: Use performance data to set cloud VM sizes. Do not replicate on-premises provisioning patterns.Modernize in waves: Once lift-and-shift workloads are stable in the cloud, identify the top candidates for modernization based on business value and team readiness Closing Thoughts There is no universally correct answer between lift-and-shift and modernize. The decision is contextual, and applying the wrong strategy to a workload modernizing something that should have been retired, or lifting-and-shifting something that needed to be rebuilt creates costs that compound over time. The framework above does not eliminate judgment. It structures the judgment so it is applied consistently, documented, and defensible to stakeholders who will inevitably ask why you chose the path you did.

By Srinivasarao Thumala

Monthly Top Methodologies Experts

expert thumbnail

Stefan Wolpers

Agile Coach,
Berlin Product People GmbH

AI for Agile Coach, Scrum Trainer with Scrum.org. Author of the “Scrum Anti-Patterns Guide.”
expert thumbnail

Oreoluwa Omoike

DevOps Engineer,
Procoreplus

Oreoluwa is a Site Reliability Engineer with a strong interest in DevOps, cloud platforms, and applied AI. She focuses on building scalable, efficient systems and enjoys exploring how intelligent tooling can improve reliability and operations.

The Latest Methodologies Topics

article thumbnail
AI on Top of a Dysfunctional System
Explore 10 Product Backlog anti-patterns that AI can make worse by masking dysfunction, missing evidence, and flawed decision-making.
October 2, 2026
by Stefan Wolpers DZone Core CORE
· 425 Views
article thumbnail
From Wild West to Context-Driven Engineering: How Culture Shapes Software Decisions
Engineering culture shapes architecture. Explore Wild West, Bureaucratic, and Context-Driven models for balancing governance, autonomy, accountability, and AI.
October 2, 2026
by Otavio Santana DZone Core CORE
· 394 Views
article thumbnail
Can Your Team Name the Work It Already Runs With AI?
The AI Workflow Inventory gives teams a shared list of recurring AI work, creating the foundation for A3 delegation decisions.
September 25, 2026
by Stefan Wolpers DZone Core CORE
· 1,316 Views · 2 Likes
article thumbnail
Three Hidden Traps That Shape Software Engineering Decisions
Learn how cognitive biases like status quo, complexity, and broken windows affect software engineering decisions, code quality, and technical debt.
September 23, 2026
by Otavio Santana DZone Core CORE
· 1,679 Views
article thumbnail
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures
AI is generating code faster than humans can review it. The fix is cognitive architectures that understand not just "what changed" but "why" and whether it's safe.
September 22, 2026
by Sayan Chatterjee
· 2,110 Views · 2 Likes
article thumbnail
Architecting Production AI Across Clouds: Patterns That Decide System Survival
In production, enterprise AI rarely fails at the model. It fails in the architecture around it. Here are the cross-cutting patterns that work.
September 16, 2026
by VenkataSrinivas Kantamneni
· 2,815 Views
article thumbnail
Alert Fatigue as a System Design Problem: Engineering On-Call Reliability in Modern SRE Teams
Alert fatigue from excessive notifications exhausts on-call engineers, eroding SRE culture. True reliability requires resilient system design, not heroic human effort.
August 21, 2026
by Oreoluwa Omoike
· 1,676 Views
article thumbnail
Reliability Without Control: Operating SRE Practices in Platform–SaaS and API-Dependent Systems
Modern SRE shifts focus from component health to user experience, relying on accurate signals and human response to sustain reliability despite reduced control.
August 20, 2026
by Oreoluwa Omoike
· 1,746 Views · 2 Likes
article thumbnail
From Agile to the Product Operating Model
Survey Results: From Agile to the Product Operating Model — learn what practitioners say is actually changing in product management, teams, and delivery now.
August 17, 2026
by Stefan Wolpers DZone Core CORE
· 1,303 Views · 1 Like
article thumbnail
Code Generation Is Solved; Trust Is the Bottleneck
Polygraph is an open-source Claude Code plugin that finds bugs in stateful code (reducers, workflows, checkout flows, session managers) in any language.
August 14, 2026
by Jean-Jacques Dubray
· 2,346 Views · 1 Like
article thumbnail
Incident Management and the Rise of AI SRE Agents
A newer category, dedicated AI SRE agents, goes further: they actively query logs, metrics, and deploy history live during an incident.
August 11, 2026
by Vidyasagar (Sarath Chandra) Machupalli FBCS DZone Core CORE
· 2,603 Views · 2 Likes
article thumbnail
I Got Tired of Copy-Pasting Microfrontend Boilerplate, So I Built a Bridge
A tiny, type-safe bridge for mounting React microfrontends across Module Federation boundaries — without repetitive lifecycle wrappers, shared stores, or code generation.
August 10, 2026
by Vitaly Zheltko
· 3,935 Views · 1 Like
article thumbnail
Building Internal Developer Platforms as Products: A Practical Guide for IDP Architects
Successful IDPs aren't built on technology alone — they combine platform engineering with product thinking and developer experience.
August 7, 2026
by Josephine Eskaline Joyce, Ph.D DZone Core CORE
· 2,018 Views · 2 Likes
article thumbnail
AI in SRE: A Practical Autonomy Model for Self-Healing Infrastructure
A practical framework for graduated autonomy in self-healing infrastructure, covering three remediation tiers and policy-driven blast-radius controls for cloud SRE teams.
July 29, 2026
by Shraddhaben Gajjar
· 3,485 Views · 2 Likes
article thumbnail
Lift-and-Shift vs. Modernize: A Decision Framework for Enterprise Workloads
This five-question framework helps you decide for each workload whether to rehost, re-platform, or rebuild before you migrate.
July 24, 2026
by Srinivasarao Thumala
· 4,132 Views · 1 Like
article thumbnail
One of Waterfall's Most Resilient Artifacts
Learn about why QA-as-a-phase persists, the costs it creates, and how teams can transition to continuous quality across the software development lifecycle.
July 24, 2026
by Stelios Manioudakis DZone Core CORE
· 4,011 Views · 3 Likes
article thumbnail
The Rise of Agentic SRE: Humans, Agents, and Reliability
Agentic SRE speeds up incident response, but it also requires clear guardrails, strong observability, and human oversight.
July 23, 2026
by Neel Shah
· 4,929 Views · 1 Like
article thumbnail
Spec-Driven Development Renamed an Old Problem; It Didn't Solve It
Spec-driven development improves AI coding, but keeping specs in sync remains the same challenge teams have faced with docs and READMEs for years.
July 21, 2026
by Sam K
· 4,095 Views · 1 Like
article thumbnail
Agent Sprawl Is Your Next Production Incident: An SRE Response to Datadog's State of AI Engineering 2026
Datadog published the State of AI Engineering 2026 report. Read it. It's the most comprehensive look at AI in production available now.
July 20, 2026
by AJAY DEVINENI
· 4,079 Views · 1 Like
article thumbnail
7 Essential Guardrails for Building AI SRE Agents
AI agents can take over the first minutes of incident response, but only with the right boundaries. Seven guardrails that keep an SRE agent from becoming the outage.
July 20, 2026
by Akhilesh Rao Meesala
· 2,697 Views · 3 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×