The Agile methodology is a project management approach that breaks larger projects into several phases. It is a process of planning, executing, and evaluating with stakeholders. Our resources provide information on processes and tools, documentation, customer collaboration, and adjustments to make when planning meetings.
AI on Top of a Dysfunctional System
Can Your Team Name the Work It Already Runs With AI?
TL; DR: The AI Delegation Lifecycle Your team ships AI outputs that nobody fully trusts; you needed to be quick, and “dirty” tagged along. Too bad that ungoverned automation becomes AI debt when a stakeholder asks who owns it. But do not despair: The AI Delegation Lifecycle turns skills you already use into six decisions you can apply this week to govern that work and prove it audit-ready and suited for agent harnesses. Popular Traps When Creating AI Output All teams can show you what their AI produces: status reports sent without anyone touching them, release notes drafted in seconds, a customer-facing FAQ that updates itself. Far fewer teams can answer the question posed by a prospective customer or by compliance: how do you govern your own internal AI use? Often, in the frenzy past to make of AI, nobody decided. Outputs without decisions are expensive, as nobody: Decided that the status report should run unattendedWrote down what a good output looks like,Analyzed the effects of a recent model change, orChecked last month whether it still produces one. The work grew that way, one helpful shortcut at a time, until it became a system nobody could explain, and nobody owned. As we know, complex systems always start as complicated systems that, at least to some people, still seem understandable. That approach of avoiding the creation of that gap, or “evolution”, is to address one decision at a time systematically. I call the practice the Delegation Lifecycle, and you already have most of the skills it requires. How Outputs Pile Up Without a Single Decision Behind Them The mechanism is ordinary yet still a shortcut: a Scrum Master pastes Retrospective data into ChatGPT to save 20 minutes. Next Sprint, the team does it again, as it worked the first time. By the third month, the summary goes straight into the team wiki, and nobody reads the raw notes anymore. The team charter does not cover it, but it is convenient; “accepted” as an amendment to the working agreement by not opposing it. The shortcut became infrastructure while everyone was busy delivering. Now multiply that by every person on the team and every task an AI can touch. You get what I call AI debt: a pile of useful, undocumented, unowned automation that works right up until the moment someone asks who is responsible for it. The “we ship it now and fix it later” habit that got you through the funding round, the launch, the reorg, and the last crisis in general, becomes the liability that shows up the moment you least expect it. The problem is not that the team uses AI, but that it uses AI without deciding which decisions belong to it. Agile practitioners are good at making decisions. We make working agreements, we set acceptance criteria, we run Retrospectives, and we refine backlogs. The Delegation Lifecycle takes those habits and points them at the work you have started handing to a model. What Counts as Delegation in the Delegation Lifecycle First, a boundary, because not every AI interaction needs governing. By delegated work, I mean recurring work in which an AI output is incorporated into a team artifact, stakeholder communication, operational workflow, or customer-facing surface. Asking a model for five ideas before you write the update yourself is not delegation. That is assistance, and it needs prompt discipline, a skill, but not a governance record. The model drafting the update from your tracker every Friday is a form of delegation. And so is sending it without a human in the loop. The lifecycle applies when AI use becomes recurring, consequential, externally visible, or embedded in how the work runs. Not every prompt needs a record, but every delegated workflow does. One Decision, Six Stages Take a single piece of work your team has delegated to AI. Not the whole AI strategy. One task: the status report, the test generation, and the first-draft release notes. That one decision has a complete path, from “should AI touch this at all” to “what do our records prove when someone asks?” The Delegation Lifecycle places a decision at each of the six points along that path. Decide Question it answers: Should AI touch this work, and how autonomously?Agile skill you already have: Decision-making with explicit categories. Route and Boundaries Question it answers: Which model, tool, environment, data boundary, and sufficiency bar are appropriate?Agile skill you already have: Acceptance criteria and process design. Hand Over Question it answers: How does the work transfer, and who owns it?Agile skill you already have: Working agreements. Define Done Question it answers: What must be true before the output leaves the team, and how is it evaluated?Agile skill you already have: Definition of Done. Inspect Question it answers: Is the delegation still safe and useful?Agile skill you already have: Retrospective facilitation. Roll Up Question it answers: What evidence can we show the people who ask?Agile skill you already have: Stakeholder communication. Before the first stage, one rule sets up the rest. The A3 Framework is the entry gate: Assist when AI supports human judgment, and you own the outcome,Automate when AI executes a bounded task under human-owned responsibility, andAvoid when the work is too consequential, ambiguous, or sensitive to hand over. If you followed my work here on the blog or took the earlier version of my AI4Agile course, you already use it. A3 decides whether work enters the lifecycle at all; the six stages sketched above govern what happens once it does. Stage 1, Decide: Should AI do this work, and at what level of autonomy? The skill underneath is decision-making with explicit categories. Where teams get stuck: Assist work quietly becomes Automate work when the review habit disappears. You start by checking every output, then most of them, then none, and nobody decided that on purpose. Stage 2, Route and Boundaries: Which model, tool, and environment run this, what data may enter, and what counts as good enough? Not every task needs your most expensive model, and not every task can run on your cheapest. Model routing means designing the process and acceptance criteria for model, tool, and data choices. The skill is the one you use every time you define done for a Product Backlog item. Where teams get stuck: they default to the priciest model for everything, never set a sufficiency bar, and then cannot explain the monthly bill. A cost nobody can account for — better: a low return on invested tokens — is a Stage 2 failure. Stage 3, Hand Over: How does the work transfer, and who owns it? Task split, owner, inputs, outputs, validation, stop rules, and a record: this is a working agreement, written for a collaborator who happens to be a model. The skill is the same one you use to set team norms. Where teams get stuck: the handoff lives in one person’s head, and no human owner is named. When that person changes teams or leaves, the system leaves with them. Without stop rules, nothing halts the work when the output starts drifting. Stage 4, Define Done: What must an AI-assisted output meet before it leaves the team? Stage 3 is how the work transfers; Stage 4 is the release gate it has to pass: the verification level, provenance disclosure, data hygiene, and the sufficiency tier from Stage 2. This is your Definition of Done, extended to work that a model touched. Where teams get stuck: “looks good” becomes the only “standard.” Approval gets mistaken for review. Someone clicks send on a procurement email the model wrote after skimming it, and now the team’s name is on a claim nobody verified. Approval is not review, and once an external audience is involved, that gap is what throws your team under the proverbial bus. Stage 5, Inspect: Is the delegation still working, or has it drifted? This is a Retrospective focused on delegated work rather than the team: Has the output quality slipped?Has Assist crept into Automate? Admittedly, applying “evals” sounds fancier, but the skill is Retrospective facilitation, which you run every Sprint. Inspection does not mean reviewing every output forever. It means agreeing on a sampling rate, the drift signals worth watching, and the trigger that brings the work back under tighter human review. Where teams get stuck is a set-and-forget mentality: Nobody scheduled the inspection, so the drift compounds unseen until it becomes an incident. Stage 6, Roll Up: What does all of this prove to the people who ask? Leadership, enterprise procurement, and increasingly, regulators want evidence of controlled AI adoption. The skill is stakeholder communication. This stage needs no separate governance artifact. The records from Stages 1 through 5 should already aggregate into what those people ask for: a delegation inventory, an autonomy distribution across Assist and Automate, and an inspection trail. Where teams get stuck is in governance theater. They build a separate leadership deck full of confident claims, disconnected from what the team actually does, and a single sharp question from a CFO collapses it. The Stages of the Delegation Lifecycle Are a Loop, Not a Checklist The six stages are a teaching order, not a strict sequence. In practice, your team will agree on the Definition of Done while filling out the A3 handoff canvas, and a finding from an inspection will send a task straight back to Stage 1 for re-classification. That is the system working, not failing. The stages also depend on each other. A Stage 1 decision that never gets inspected becomes the most dangerous kind of automation: confident, unattended, and unowned. A Definition of Done with no handoff behind it has no teeth, because nobody agreed on who applies the standard or when. Just count how many of these six stages your team has a real decision behind right now; it will make a good starting point for a team discussion on how you are using AI at the moment. Two points on the path are deliberately set to have no artifact. Before Stage 1 sits, know which work your team does, at what frequency, and at what stakes, which is a forensic analysis of your own workflow. Around all six sits your AI working agreement, the team norm layer. Neither needs a new canvas. The AI Delegation Lifecycle adds a document only where a recurring decision genuinely had no home. Why This Is Not Optional Anymore The approach the Delegation Lifecycle proposes is not just about internal hygiene. AI use is moving from personal productivity into organizational accountability. Since February 2, 2025, Article 4 of the EU AI Act has required providers and deployers to ensure a sufficient level of AI literacy among staff and others operating AI systems on their behalf. (Which, interestingly, seems to be largely ignored by many players.) Enforcement through national market surveillance authorities will take effect on August 3, 2026. NIST organizes AI risk management around the four steps: govern, map, measure, and manage. Anthropic’s first Economic Index found that real-world Claude usage already splits between augmentation and automation: 57% augmentation, 43% automation. The practical question underneath it all is simpler: can you show the decisions behind the work you delegated? The Roll-Up Is the Quiet Payoff Most teams miss the byproduct of every stage, producing a valuable record as part of normal use: The A3 decisions become a portfolio of what you deliberately automated, assisted, and kept human.The routing records become AI spend by task and tier, with a reason attached.The Definition of Done sign-offs and inspection logs serve as an audit trail of controlled, inspected adoption. Nobody fills in an extra report: operational work generates governance evidence as it runs. So when a prospect asks, “How do you govern your own internal AI use?” the team running this lifecycle does not shrug. It answers with its records. That is the difference between a team that merely uses AI and a team that can be trusted with it, and that trust is becoming a line item in enterprise procurement. Your Turn Pick one task your team has handed to an AI. Just one. Walk it through the six questions out loud in your next Retrospective: did we decide this, who owns it, what does done mean, when did we last check it, and what would we show someone who asked. I guess that you will find at least one stage where the honest answer is “nobody decided that.” That is your starting point: adopt one stage at a time, wherever your pain is sharpest. Knowledge that walked out with a departing colleague points to Stages 3 and 4. A token bill nobody can explain to the CFO points to Stage 2. An output that embarrassed you in front of a stakeholder points to Stages 4 and 5. Conclusion I turned this lifecycle into a working method teams can apply immediately: how to decide what to delegate, hand it over safely, inspect for drift, and produce evidence without building a governance theater. Which of the six stages does your team actually have a decision for? I am curious, and I suspect the answer is fewer than you would like.
TL; DR: The “Agile to the Product Operating Model” Survey Results Between August 2 and August 10, 2026, 48 practitioners participated in my Agile to Product Operating Model (POM) survey, which tries to shed light on what is actually changing. Let me summarize the answers for you: the reported transformations change decision-making less than the Cagan framework suggests. Where respondents report improvements, they appear in delivery and collaboration rather than in business results. Unfortunately, the human side of the transition is the least encouraging part of the answers. Thesis: Product operating model transformations mostly change vocabulary and organizational structure while leaving the decision system — who decides what gets built, on what evidence, at what speed- largely untouched. AI is changing product decisions independently of POM transformations. Who Answered, and Why I Report Counts, Not Percentages Of the 48 respondents of the Agile to Product Operating Model, 26 work in organizations that have adopted a product operating model or are moving toward it (we refer to them as the “movers” from here on). The other 22 work in organizations that are not moving; these organizations are still discussing the issue, have decided against it, or have never considered it. So you get absolute counts, and every finding below is directional, not definitive. Additional limitations are: 24 of the 48 respondents are Scrum Masters or Agile Coaches, the roles under the most pressure from this shift.38 of 48 work in organizations with more than 250 people; the sample contains 2 respondents from startups and 2 from scale-ups. Apparently, the people who have participated in the survey are those who still care enough about agile practice to read a newsletter about it; people who left “Agile” entirely are structurally missing. Also, as I did not ask respondents to identify employers, the 48 participants report on their organizations, not 48 distinct verified organizations. Putting it all together, this is a report from inside large organizations, where the product operating model makes its boldest promises. One Fundamental Change in 26 Answers Asked which statement best reflects their experience so far, the 26 movers split like this: 9 say mostly relabeling, the same way of working with new vocabulary9 say real change in some areas, relabeling in others7 say it is too early to tell. Exactly one participant reports a fundamental change in how the organization decides what to build, which is the most interesting learning from the responses: changing the vocabulary, changing the organizational structure, and changing how decisions actually get made are three different things. Better Delivery, Inconclusive Business Results, Worse Morale The Agile to Product Operating Model survey asked movers to compare six outcomes against their previous way of working. Aggregating “somewhat” and “clearly” in each direction: Speed of delivery: Better (9), No change (10), Worse (1), Too early / cannot tell (6)Value delivered to customers: Better (8), No change (8), Worse (2), Too early / cannot tell (8)Business results: Better (3), No change (8), Worse (4), Too early / cannot tell (11)Team motivation and morale: Better (7), No change (4), Worse (9), Too early / cannot tell (6)Developers’ satisfaction: Better (4), No change (8), Worse (7), Too early / cannot tell (7)Collaboration with stakeholders: Better (9), No change (6), Worse (4), Too early / cannot tell (7) The table does not say that the transitions are failing: Delivery speed leans positive: 9 better against 1 worse.Customer value leans positive: 8 against 2.Collaboration with stakeholders leans positive: 9 against 4.Business results are inconclusive, with 11 of 26 unable to say and the rest split. The people dimensions are the only ones that lean consistently negative: morale at 9 orgs is worse than at 7, where it improved; developers’ satisfaction is similar: worse at 7 orgs and better at 4 orgs. Two ways to interpret this information, and the survey does not help to choose: (1) The improvements concentrate on the things enterprises already know how to optimize (flow, coordination, delivery), while the transformations may be falling short of the things the model claims to change most fundamentally. (2) Operational effects precede commercial effects, because business outcomes have longer feedback cycles, and 11 of 26 respondents explicitly cannot yet say. Consider that flow, coordination, or delivery are easier to measure and attribute than commercial effects, which is a significantly fuzzier area. Therefore, the interesting question is how this changes over time, not how it looks at one moment. The range of individual experience behind those aggregates is wide. One respondent, 12 months into adoption at a large financial-services organization, described the hardest part as “surviving the political game that the POM shift has been” and reports worse answers on five of the six dimensions. The other two respondents who have been at 12 or more months report the opposite: better value, better morale, and better collaboration. Three long-tenured accounts cannot agree on whether the practical transition mechanics, culture, affected products and services, or the organization itself makes the difference. (That is the curse of a small sample size.) Agile to Product Operating Model: The Empowerment Gap, and What Settles Uncertainty Cagan’s own test separates empowered product teams from feature teams: in a product team, the team is tasked with solving a problem and owns the solution. In a feature team, “the value and business viability are the responsibility of the stakeholder or executive that requested the feature.” The survey asked movers which pattern operates today, in the middle of their transitions: 9 of 26 report the empowered pattern (leadership sets goals and problems to solve, teams decide what to build)8 of 26 report that leadership decides which features get built and teams implement them.4 of 26 report stakeholder requests are filling the product backlog reactively.4 of 26 say it is genuinely unclear or contested right nowA single one reports teams setting their own direction. So even among respondents whose organizations are actively adopting a model whose entire premise is empowered teams, the “feature-list or reactive-backlog” pattern (12) outnumbers the “empowered” pattern (9). A related question asked what settles it when the organization is uncertain whether something is worth building: 6 said debate and prioritization before anything gets built5 said whoever has the most authority decides5 said things just get built and shipped, and an explicit decision rarely follows. 4 said research evidence5 said it varies too much to say. In total, only five respondents fall into the categories I pre-registered as evidence-led: research evidence, disposable prototypes, or ship-and-measure. There is one case among all the answers: a respondent reports that a disposable AI-assisted prototype reduces build uncertainty: build to learn, decide, and throw it away. That same respondent, 12 months into adoption, also reports better responses across all six outcome dimensions listed above. Of course, one case is not evidence of a relationship. It is a question worth asking in a larger sample, not a claim I am entitled to make here. In Little Code, Big Waste, I argued that cheap AI code removes the cost gate that used to force a should-we-build-this decision, and that “when generating plausible code becomes cheap, every hour spent building the wrong thing becomes waste that can now be produced at scale”. This survey says the prototype-to-decide pattern barely exists yet in this sample. Agile to Product Operating Model Survey Shows: The Two Transitions Run Separately If you believe the keynote version of events, AI is forcing organizations to redesign their operating model around it. The 26 movers report something different. Only one respondent says AI is a core reason for the change; 6 of 26 call it one factor among several; 16 say AI adoption runs in parallel, but separately from, the operating model change; and 3 report little or no role. Now the other side. Among the 22 respondents in non-moving organizations, 15 report that AI is changing how product decisions are made anyway: 6 noticeably, 9 in pockets, without any formal model change. Despite the small sample size, the data support a narrow claim: for most movers in this sample, AI and operating-model transformation are separate initiatives, while for most respondents in non-moving organizations, AI is changing product decisions without any model change at all. Formal operating-model change is neither a prerequisite for AI-driven changes to product decision-making nor, so far, organized around them. The survey did not measure how AI enters these organizations or who authorizes it, so I will not claim more than that. The disconnect between the two is already the finding. My Pre-Survey Hypothesis Scoreboard Here are my verdicts on my four hypotheses, H1-H4, that led to the creation of the Agile to Product Operating Model survey: H1, consultant-driven transitions produce more relabeling than leadership-driven ones: Relatively consistent, but no verdict: 3 of the 4 respondents whose transition is driven by consultants or a transformation office report “mostly relabeling,” against 4 of 16 in product-led or executive-led transitions. But those are just four participants. H2, the empowerment gap: The feature-list plus reactive-backlog patterns (12) outnumber the empowered pattern (9). That is observed in this sample, however, not established beyond it. H3, where AI drives the move, product judgment is the scarcest capability: incorrect on my side: Among the 7 respondents whose organizations treat AI as a core reason or a factor in the move, the most frequent scarcity answer is “delivery capacity is still the bottleneck” (3), ahead of product judgment (2). Across all 26 movers, the scarcity question splits three ways: “stakeholder alignment and decision speed” (7), delivery capacity (6), and product judgment (6). One distinction keeps H3’s rejection from closing the question. The survey measures perceived constraints, and perception rarely reflects reality accurately. AI may have changed the economics of implementation faster than organizations have updated their sense of where the workflow bottleneck sits. Whether delivery capacity objectively remains the constraint is a different question, and this survey cannot answer it. What it can say: practitioners today experience decision speed, delivery, and judgment as roughly competing constraints, and the judgment-scarcity era I anticipated is not the world they report living in. H4, organizations that build without explicit decisions report worse customer value than evidence-led ones: rejected as stated: 1 of 5 build-without-deciding respondents reports better customer value, against 2 of 5 evidence-led ones. Again, the number of replies is too small. Changing the Structure Without Changing the Decision System Three of the findings above belong side by side. Only 1 of 26 movers reports a fundamental change in how build decisions are made. Only 9 of 26 report the empowered decision pattern operating today. Only 5 of 26 fall into the evidence-led categories for resolving build uncertainty. Together, they suggest a more precise diagnosis than “the transformation is theater.” Transformations change three different layers, and the layers move at different speeds: Vocabulary changes first: Product operating model, empowered teams, outcomes over outputs; you get the idea.Structure changes second, and often really does change roles, reporting lines, team topologies, or artifacts.The decision system changes last, if at all: Who decides, based on what evidence, under what uncertainty, with what authority, at what speed. This survey reads like a snapshot of organizations renaming the first layer, reorganizing the second, and leaving the third largely untouched. That is why the 1-of-26 count deserves the weight I put on it earlier. One of the product operating model’s defining promises is a different way of deciding what to build. If only one of the 26 respondents experiences a fundamental change in that mechanism, the question is no longer whether the transformation is proceeding fast enough, but what is being transformed. That mental model offers one possible explanation for the outcome table: structural change could improve coordination and flow before it alters the quality of product decisions. Whether that explains the pattern here is impossible to establish from 26 responses. A long-term study of organizations and involved practitioners would need to test whether decision-system change predicts eventual business outcomes. (Consider that a hypothesis this survey generated, not a result it delivered.) Agile to Product Operating Model: Why Product Washing Is Easier Than Empowerment The survey cannot tell us why the decision system resists change. What follows is my hypothesis, argued from two decades of watching transformations, not a survey result. The tempting explanation is that these organizations are implementing the product operating model badly, and that a proper implementation would deliver. I spent those two decades watching the Agile community run exactly that defense. Every failed adoption was “not real Scrum.” The argument is unfalsifiable, and it taught an entire industry to blame practitioners instead of examining incentives. I will not run the same defense for the product operating model. In November 2024, I described Product Washing: the hollow adoption of product practices that “leaves companies stuck in the same old dynamics but with a new vocabulary,” transformation by reprinting business cards. My hypothesis for the mechanism: product washing is not an implementation failure but what enterprise incentives produce when you ask powerful people to redistribute their own power. The model demands that stakeholders with budget authority hand problem-selection to product leadership and solution-selection to teams. Budget authority is power, and in many large organizations, “product leadership” has become a new title for the same stakeholders who control the money. Meanwhile, middle layers face asymmetric career payoffs: a visible failure damages a career far more than a shared success advances it. Under those payoffs, routing decisions through committees and sign-offs is rational self-protection, and it does not vanish because the org chart was redrawn. Two honest caveats bound this hypothesis: First, in regulated industries, some of that scrutiny is very relevant: when a named person must answer to a regulator, a sign-off chain is accountability, and the respondents from banking and the public sector live with constraints no product operating model erases. The skill worth having is telling required governance apart from accountability theater; most organizations run both and label neither. Second, this survey is a cross-section, not a time series, and two competing explanations fit the same data: The transition hypothesis: decision authority changes slowest of all the layers, and these organizations are simply not there yet.The attractor hypothesis: enterprise incentives pull transformations toward renamed feature factories and hold them there. My incentive argument predicts the attractor. With only 3 respondents at 12 or more months, this survey cannot distinguish between them. That is a testable question for a future survey. The Non-Movers’ Catch-22 The 22 respondents in non-moving organizations deserve more attention than transformation literature usually grants them. Their top reasons for staying put: Leadership sees no need or has other priorities (15),Lacking the product-management maturity to build on (13), andThe cost and fatigue of yet another transformation (10). Only 6 claim their current way of working performs well enough. The dominant non-mover position is not a confident endorsement of the status quo. And inside the reasons sits a genuine Catch-22. Organizations supposedly need the product operating model because their product capabilities are weak, while 13 of 22 respondents say weak product capabilities are precisely why their organization cannot adopt it. A transformation that requires the maturity it promises to create is a hard sell to people who have already survived several that made the same offer. Conclusion: Three Questions for Monday Morning Three questions locate your own organization on this map: First: who decided the last thing your team built, and would your CPO name the same person? If the answers differ, you have found the gap between the model on the slides and the model in operation. Second: what settled your organization’s last genuinely contested build decision: evidence, debate, or seniority? “Whoever has the most authority decides” got 5 votes out of 26 in this survey. Be honest about whether your organization would add a sixth. Third: what would have to be true for a disposable prototype to settle the next contested decision instead? That is a political question, not a technical one: who would have to accept evidence as a tiebreaker, and what would it cost them? Forty-eight answers later, the better question is no longer which operating model organizations adopt. It is how they make decisions when the cost of trying something has collapsed while the cost of deciding has not: Who decides on what evidence, and how quickly can the organization act on what it learns? That is where my work is heading next.
Picture this: you are a software developer building an education platform, and you receive from the product owner some requirements written in business language (Gherkin). You need to implement these scenarios in Python. Probably you will start creating models and service modules. You will create some classes to represent the entities described in the scenarios, like Student, Course, and Subject. You will add conditionals and loops in the entity classes to control the business logic and restrict paths in the code: Python # Enroll a student in a course if course.status == "active" and student.course == None: student.course = course raise BusinessError("Student already in a course") Also, you will create a class to represent the persistence layer (database) and methods like list_students, get_course_by_name, and create_student to add, delete, update, and return data from the database. You will probably create facades to group the classes in a logical sequence, add more ifs, elses, and loops to control the code flow. At the end of the sprint, you have a scenario implemented and tested. There is nothing wrong with its style of implementation. It is a common process. However, something loses importance in this process: the business scenario itself. In this article, I’ll showcase a behavior-driven development approach that converts business languages directly to executable code. The intention is to keep the implementation closer to the business language and promote the code to the source of truth. Gherkin Scenarios for the Education Platform Going back to the fictional (not much) story. Here are some scenarios an education platform could have: Gherkin Feature: Student GPA and approval Scenario: Student is approved when GPA is 7 or higher and all subjects are passed Given a student named "John" is enrolled in the "Computer Science" course And the course has the subjects "Math", "Physics", and "Programming" And the student has the following grades: | Subject | Grade | | Math | 7 | | Physics | 8 | | Programming | 9 | When the system calculates the student's GPA Then the GPA should be 8 And the student status should be "Approved" Feature: Student enrollment in course subjects Scenario: Student cannot enroll in a subject from another course Given a student named "Carlos" is enrolled in the "Medicine" course And the subject "Algorithms" belongs to the "Computer Science" course When the student tries to enroll in the subject "Algorithms" Then the enrollment should be rejected And the system should show the message "Students can only enroll in subjects from their own course" Feature: Student enrollment in a course Scenario: Student enrolls in an active course Given a course named "Architecture" is active When a student named "Julia" tries to enroll in the "Architecture" course Then the enrollment should be accepted And the system should show the message "Student enrolled in course" Feature: Course cancellation Scenario: Students cannot enroll in a canceled course Given a course named "Architecture" has been canceled by the general coordinator When a student named "Julia" tries to enroll in the "Architecture" course Then the enrollment should be rejected And the system should show the message "Canceled courses cannot accept new enrollments" They are pretty, readable, easy to understand, and find inconsistencies. Now, a possible implementation as described in the previous story. It was simplified for the sake of this article. Let us look at a more traditional implementation. Python # Entities class Course: def __init__(self, course_id, name): self.course_id = course_id self.name = name self.is_canceled = False class Student: def __init__(self, student_id, name): self.student_id = student_id self.name = name self.course = None # Application class UniversityService: def __init__(self): self.courses = {} self.students = {} def create_course(self, course_id, name): self.courses[course_id] = Course(course_id, name) def create_student(self, student_id, name): self.students[student_id] = Student(student_id, name) def cancel_course(self, course_id): course = self.courses.get(course_id) if course is None: raise ValueError("Course not found") course.is_canceled = True def enroll_student_in_course(self, student_id, course_id): student = self.students.get(student_id) course = self.courses.get(course_id) if student is None: raise ValueError("Student not found") if course is None: raise ValueError("Course not found") if course.is_canceled: raise ValueError("Canceled courses cannot accept new enrollments") student.course = course # Scenario: Students cannot enroll in a canceled course service = UniversityService() service.create_course("C1", "Architecture") service.create_student("S1", "Julia") service.cancel_course("C1") try: service.enroll_student_in_course("S1", "C1") print("Unexpected result: student enrolled in a canceled course") except ValueError as e: print(e) It was done in a traditional style. Notice the technical references like service and the preconditions and business logic spread in many ifs in the code. We forgot to represent the system behavior in a simple and explicit way. The scenario was spread into many pieces, and it may be hard to put all of them together when we need to understand the code in the future. Consider that more features will be integrated into the code, and more if/else statements will be introduced to control the business logic and new flows. In summary, the scenario cannot be read as it was presented by the business team. It is hard to validate that the system is doing what it should do without proper unit tests and careful code review. We can try to test its integration with Python Behave to bring the explicit behavior back to the game, but it may be hard to do it without coming up against technical stuff like services. The system works, but it is hard to prove that it behaves as expected just by reading the code. At this point, the development team and the business team are not talking the same language anymore. There is a translation from business language to production code (technical stuff). Behavior-Driven Development Now, using the framework Guará to represent the scenarios directly in the code. The code now tells the story. For example, the scenario Student enrollment in a course can be written like this: Python from guara.application import Application eduapp = Application() ( eduapp.given(IsActiveCourse, course_id=course_id) .and_(IsNotStudentInACouse, student_id=student_id) .when( EnrollStudentInCourse, student_id=student_id, course_id=course_id, ) .then(it.IsEqualTo, "Student enrolled in course") ) The preconditions IsActiveCourse and IsNotStudentInACourse are now explicit and are at a higher level of the code. Not buried in the methods in the form of if conditionals. The precondition and action classes have single responsibilities. Python from guara.transaction import AbstractTransaction class IsActiveCourse(AbstractTransaction): def do(self, course_id): print(f"Checking the status of course {course_id}") status = database.courses.get_status(course_id=course_id) if status == "Active": return True raise CourseCanceledException("Course canceled") class IsNotStudentInACourse(AbstractTransaction): def do(self, student_id): print(f"Checking if student in a course") course = database.student.get_course() if course: raise StudentException("Student already in a course") class EnrollStudentInCourse(AbstractTransaction): def do(self, student_id, course_id): print(f"Enrolling student {student_id} in course {course_id}") status = database.enroll_course(course_id, student_id) return "Student enrolled in course" In the end, it is easier to compare the code against the scenario steps and assert they are present in the code. Python import argparse from guara.transaction import Application from guara import it eduapp = Application() def main(): parser = argparse.ArgumentParser() parser.add_argument("--action", required=True) parser.add_argument("--student-id") parser.add_argument("--course-id") args = parser.parse_args() if args.action == "enroll_course": try: ( eduapp.given(HasCourse, course_id=args.course_id) .and_(IsActiveCourse, course_id=args.course_id) .and_(HasStudent, student_id=args.student_id) .and_(IsNotStudentEnrolledInCourse, student_id=args.student_id) .when( EnrollStudentInCourse, student_id=args.student_id, course_id=args.course_id, ) .asserts(it.IsTrue) ) except Exception as e: print(str(e)) app.undo() # Calling the CLI python edu.py enroll-course --course-id 10 --student-id 1324 Benefits The production code is now the source of truthIt can be compared directly to the business scenariosThe responsibilities are encapsulated in dedicated classesIt is possible to undo operations easily once the framework is based on the Command Pattern (GoF)It is easy to add more behavior to the code without changing other classesThe classes are reusableIt hides the technical stuff. They still exist, but now the actions are first-class citizens Points of attention It is not a one-size-fits-all style. It is necessary to evaluate whether the system under development will benefit from this code styleMakes more sense when the scenarios are defined in Gherkin language; otherwise, it will be necessary to translate the requirement to code as done in the traditional implementation Conclusion The important difference is that the source code still reads almost like the original Gherkin scenario. Instead of hiding business rules inside technical layers, we keep them visible and explicit in the code.
In traditional software delivery, requirements often functioned as an initial alignment artifact. Once development began, implementation, iteration, and testing took precedence. AI-assisted development changes that balance. As AI systems generate increasing portions of code, tests, and documentation, the quality of input becomes the primary determinant of output quality. Requirements no longer describe work. They define the execution context. This introduces a structural shift. AI systems do not interpret intent independently. They operate on the context they receive. When requirements are incomplete, ambiguous, or disconnected from system constraints, the generated output reflects those limitations. This results in plausible implementations that may satisfy surface-level expectations while diverging from business or architectural intent. Requirements as Executable Context Requirements are evolving from static documentation into structured inputs for automated systems. In AI-assisted workflows, they function as the source of truth that guides generation across multiple stages of delivery. This includes: Functional intentArchitectural constraintsIntegration patternsEdge cases and non-functional requirements The more explicitly these elements are defined, the more reliable the generated output becomes. In practice, this is often implemented through retrieval-augmented approaches, where requirements, architectural decisions, and historical context are indexed and dynamically included in AI workflows. This allows generation to operate on relevant system knowledge rather than isolated task descriptions. Impact on Delivery Systems This shift has direct implications for delivery performance. First, ambiguity propagates faster. When AI accelerates implementation, unclear requirements lead to faster misalignment rather than slower discovery. The result is increased rework, even when the initial output appears correct. Second, access to context becomes critical. Requirements must be connected to architecture documentation, historical decisions, and system knowledge. Teams increasingly rely on structured knowledge bases or vector stores to ensure that relevant context is available during generation. Third, requirements become iterative. They are continuously refined as part of the delivery process, rather than treated as a fixed starting point. Changes in requirements directly affect generated outputs, which makes versioning and traceability more important. A broader industry analysis of AI-native software engineering trends points to the same pattern: delivery performance increasingly depends on structured input context, lifecycle discipline, and operating model maturity rather than generation speed alone. Practical Example: Requirement Ambiguity at Scale Consider a common scenario in AI-assisted development. A team defines a requirement for an API endpoint: “Return user account details with relevant metadata.” In a traditional workflow, ambiguity might be resolved during implementation discussions. In an AI-assisted workflow, the system may generate an implementation immediately. Without structured constraints, the generated output may: Include incomplete fieldsOmit edge cases, such as suspended accountsIgnore performance constraintsFail to align with existing API contracts The result is not always an obvious failure. The code may compile, pass basic tests, and appear correct. The misalignment becomes visible later during integration, production use, or future maintenance. This illustrates how ambiguity moves from a manageable discussion point to a source of scaled inconsistency. Practical Implications for Engineering Teams This shift requires changes in how teams approach requirements in everyday development work. Requirements need to be treated as a technical artifact, not just documentation. This typically includes: structuring requirements for machine readabilityexplicitly defining constraints and assumptionslinking requirements to architecture and system boundariesmaintaining consistency across documentation sourcesensuring traceability between requirements and generated output In AI-assisted workflows, requirements are no longer consumed only by developers. They are consumed by systems that execute based on their structure and clarity. This also affects tooling decisions. Teams integrating AI into delivery pipelines often need mechanisms to: Retrieve relevant context dynamicallyValidate generated output against requirementsTrack how requirements influence generated changesConnect requirement artifacts directly to test generation and review workflows Without these controls, AI-generated output can introduce subtle inconsistencies that are difficult to detect during review. Generated output can scale faster than a team’s ability to verify it. Implications for Engineering Practice Engineering teams need to adjust how requirements are produced and maintained. This includes: Increasing precision in requirement definitionsReducing ambiguity in edge cases and constraintsAligning requirements with system architectureUpdating requirements continuously as part of delivery The skill of defining intent becomes directly linked to delivery quality. This is particularly visible in AI-native workflows, where requirements effectively become the interface between human intent and automated execution. As AI takes on more of the mechanical work of implementation, requirements become the point where engineering judgment enters the system. They define what the model should optimize for, what it must avoid, and how reviewers should evaluate the result. Conclusion AI-assisted development does not reduce the importance of requirements. It increases it. As more of the implementation process becomes automated, the ability to define clear, structured, and context-rich requirements becomes the primary control mechanism for reliable delivery. In practice, this means that requirements are no longer a preliminary artifact. They are becoming the control layer that determines whether AI-assisted systems produce consistent, maintainable, and correct outcomes.
AI coding assistants are becoming increasingly capable at generating code, explaining systems, and accelerating development workflows. But in real engineering environments, the biggest blocker is often not the model’s ability to write code. The bigger issue is whether the assistant has the right context before it starts making changes. A developer rarely works from a single source of truth. A Jira ticket may describe the implementation task. A Google Doc may contain the detailed requirements. A slide deck may explain the business goal. A meeting summary may include key decisions, open questions, and next steps that never made it back into the ticket. For a human developer, this creates friction. For an AI coding assistant, it creates risk. The assistant may generate code that looks correct, passes basic syntax checks, and follows existing patterns - but still implements the wrong behavior because the actual feature context was fragmented across multiple places. This is where a PARA-style context workspace becomes useful. PARA - Projects, Areas, Resources, and Archives is commonly used to organize knowledge by actionability. Applied to AI-assisted software development, it can become a practical architecture pattern for preparing scattered engineering knowledge before an AI coding assistant touches code. The goal is not to dump every document into the model. The goal is to organize scattered context so the assistant can reason with the right information for the task. The Problem: AI Coding Assistants Often See Only Part of the Work Consider a developer asked to build a new data pipeline that calculates a generic quality score. The implementation sounds straightforward: Build a pipeline that joins multiple input tables, applies business rules, and produces a quality score output table. But the actual context may be spread across several sources: SourceWhat It May ContainTicketImplementation scope, acceptance criteria, due dateRequirements docBusiness rules, scoring logic, data definitionsSlide deckBusiness goal, stakeholder alignment, expected impactMeeting summaryFinal decisions, open questions, changed thresholdsExisting codePipeline patterns, naming conventions, dependency structureOlder documentsPrevious decisions, deprecated approaches, known constraints If the AI coding assistant only sees the ticket, it may miss the deeper context needed to implement the feature correctly. This is especially risky for data pipelines and analytics features, where correctness depends not only on code structure but also on interpretation: which source tables to use, how freshness should be handled, how business rules are applied, and how downstream consumers will use the output. What Can Go Wrong If the Agent Only Reads the Ticket? A ticket often captures the visible work, but not the full reasoning behind the work. If the assistant only uses the ticket, it may: Implement the task but miss business rules from the requirements documentIgnore key decisions captured in meeting summariesUse a technically available source table that is not the approved source for this featureMiss freshness expectations for the output tableProduce a score that does not match how downstream dashboards or reports will consume itFollow an outdated implementation pattern because it found old but similar codeGenerate a pull request that looks reasonable but fails product or data-quality expectations This is the core issue: The AI assistant may know how to write code, but it may not know which code should be written. That distinction matters. For coding agents to become more reliable, developers need a better way to prepare context before code generation begins. Reframing PARA for AI Coding Agents PARA can be adapted from a personal knowledge organization method into a context classification pattern for AI-assisted development. In a PARA-style context workspace: PARA CategoryEngineering MeaningAgent Context RoleProjectsActive work being deliveredCurrent feature scope, ticket, task goalAreasOngoing responsibilitiesStandards, ownership, governance, quality expectationsResourcesReusable knowledgeDocs, runbooks, design patterns, pipeline examplesArchivesCompleted or inactive knowledgeHistorical decisions, old approaches, past incidents This structure helps the AI assistant understand the role of each piece of information. A current requirement should not be treated the same way as an old design decision. A meeting decision should not be buried behind a generic document search. A reusable pipeline pattern should be available to guide implementation, while archived material should be used carefully as historical context. The value of PARA is not just an organization. It gives the assistant a way to distinguish between active task context, long-running rules, reusable references, and historical information. This flow changes how the assistant approaches implementation. Instead of asking: “What code should I generate from this ticket?” The assistant can reason from a richer question: “What is the active feature goal, what rules must be followed, what reusable references apply, and what historical context should be considered before changing code?” That shift is small, but important. Applying PARA to a Quality Score Pipeline Now apply this to the quality score pipeline example. The feature requires a pipeline that joins multiple input tables, applies business rules, and writes a quality score output table. The exact business logic is intentionally generic, but the pattern is common across analytics engineering, data engineering, machine learning platforms, and reporting systems. A PARA-style workspace could organize the context like this: Project Context This is the active feature work. It may include: The current ticketFeature scopeAcceptance criteriaCurrent implementation statusTarget output tableExpected delivery milestoneKnown blockers or open questions For the coding assistant, this answers: “What am I being asked to build right now?” Area Context This represents ongoing expectations that apply beyond this one feature. It may include: Data quality standardsFreshness expectationsOwnership rulesPrivacy or compliance constraintsNaming conventionsRelease processTesting expectations For the coding assistant, this answers: “What rules and standards must this implementation follow?” Resource Context This is reusable technical knowledge. It may include: Existing pipeline patternsSimilar transformation logicData model documentationDashboard dependency notesCommon test patternsRunbooksData validation examples For the coding assistant, this answers: “What reusable references should guide the implementation?” Archive Context This is historical information that may still be useful, but should not automatically drive the implementation. It may include: Older design decisionsDeprecated scoring logicPast pipeline migrationsPrevious quality metric experimentsHistorical meeting notesOld RCA or incident learnings For the coding assistant, this answers: “What historical context may explain why the system works this way?” The critical point is that archived context should be used for awareness, not blindly copied into the current implementation. Why Meeting Summaries Matter Meeting summaries are often underestimated in AI-assisted development. In many teams, the final decision is not always reflected immediately in the ticket or requirements document. A meeting summary may contain important details such as: A threshold was changed after stakeholder discussionA source table was rejected because of data freshness concernsA metric definition was clarifiedA downstream dashboard dependency was identifiedA launch decision was postponedAn open question was assigned to another teamA temporary workaround was approved only for the first release For a human developer, these details may be remembered from the meeting. For an AI coding assistant, they are invisible unless they are included in context. This is one reason a PARA-style workspace can be valuable. It gives meeting summaries a place in the feature context without treating them as random notes. A meeting summary tied to an active feature belongs in the Project context. A recurring decision about data freshness may become the Area context. A reusable explanation of metric calculation may become the Resource context. Once the feature is complete, the same meeting summary may eventually move into the Archive context. How the Coding Assistant Should Use Context Before Changing Code Before generating code, the AI coding assistant should use the structured context to form an implementation understanding. For a quality score pipeline, it should first understand: What the feature is trying to accomplishWhich input data sources are approvedWhich business rules define the scoreWhich decisions were finalized in meetingsWhat freshness or latency expectations existWhich existing pipeline patterns should be followedWhat downstream dashboards, reports, or consumers depend on the outputWhich historical approaches should be avoided Only after that should it propose an implementation plan or modify code. This changes the assistant’s role. It is no longer simply a code generator responding to a ticket. It becomes a context-aware engineering assistant that can reason across requirements, decisions, standards, and existing system patterns. The Bigger Shift: From Prompting to Context Preparation Prompting is still useful, but it is not enough for complex engineering work. A good prompt cannot fully compensate for missing requirements, outdated context, or scattered decisions. For AI coding assistants, the quality of the result depends heavily on the quality of the context that comes before the prompt. This is especially true when the task involves business logic, analytics definitions, data contracts, or cross-team decisions. In those cases, the question is not: “How do we write a better prompt?” The better question is: “How do we prepare the right engineering context before asking the assistant to write code?” For developers building with AI coding agents, this may become one of the most important habits: do not ask the agent to write code first. Prepare the context first. Because the future of AI-assisted development will not belong only to teams with the most powerful coding models. It will belong to teams that know how to structure knowledge so those models can make better engineering decisions.
Mainframe modernization is once again at the center of enterprise conversations. Not because something suddenly broke, but because the environment around it has changed. Organizations are being asked to move faster, integrate more easily with newer platforms, and support initiatives like cloud and AI that weren’t part of the equation a decade ago. At the same time, experienced teams are shrinking, costs are under scrutiny, and expectations from the business are higher than ever. The way organizations are approaching modernization is evolving as well. Instead of treating it as a one-time, large-scale effort, many are taking a more incremental path and making changes over time. Many are introducing more modern, agile development practices and working to bring mainframe development closer in line with how the rest of the enterprise builds and delivers code changes and manages their development cycles. Even with that shift, the same challenges still tend to surface. The Clarity Most Organizations Are Missing Most organizations approaching modernization are not lacking motivation. What’s often missing is clarity around what’s really broken, what needs to change, and what success should look like. There’s a general sense that systems are too slow, processes are inefficient, or teams are struggling to keep up. But those issues aren’t always clearly defined before decisions are made. Instead, the focus shifts quickly to solutions (new platforms, new tooling, AI) without fully understanding the root of the problem. If the issue is how work flows through the organization (how decisions are made, how teams interact, and how long it takes to move from development to production, etc.), then changing the technology alone won’t solve it. In many cases, it simply exposes the problem more quickly. Where Modernization Efforts Start to Break Down When that lack of clarity carries into execution, the gaps become much harder to ignore. Processes are often more complex than expected, approval chains are longer than they need to be, and workarounds have developed over time to compensate for inefficiencies in the official process. Introducing new tools into that environment doesn’t remove those issues; it highlights them. A faster system makes bottlenecks more obvious, and a more connected environment exposes gaps between teams. What may have been tolerated before now becomes difficult to ignore. There’s also a persistent belief in what many teams jokingly call the “magic factor.” The idea that a new platform, a new vendor, or even AI will come in and solve everything. It’s an appealing story, especially when teams are under pressure. But it sets expectations that reality can’t meet. Timelines add another layer of tension. Modernization is often scoped as a short-term project, when in reality it requires sustained effort. Training, testing, and adoption all take time, and organizations are rarely able to move as quickly as initial plans assume. Perhaps most critically, many organizations lack a true internal owner of the effort. Vendors and partners can guide the work, but they can’t drive internal adoption. When no one inside the organization is accountable for the outcome, progress slows, decisions get delayed, and momentum fades. All of this plays out against a backdrop of uncertainty. For experienced mainframe professionals, modernization can feel like a threat to years of hard-earned expertise. For newer developers, it can feel unfamiliar and difficult to navigate. Without clear communication and support, both groups can disengage. At that point, modernization doesn’t fail outright; it just never quite delivers what it promised. What Changes When It’s Done Right When organizations take a step back and approach modernization more thoughtfully, the picture can look very different. Instead of treating the mainframe as something separate, they start to bring it into the same ecosystem as the rest of their development environment. Tools like Git, modern IDEs, and CI/CD pipelines become part of the workflow. Developers no longer have to switch contexts or work in isolation. That shift alone changes how teams operate. Historically, mainframe teams have operated separately from distributed, web, and mobile teams. Each team had different tools, different workflows, and limited visibility into each other’s work. Modernization, particularly when it introduces more unified workflows, begins to break down those silos. Teams gain a clearer view of how their work connects, collaboration becomes more natural, and knowledge starts to move more freely across the organization. That has a real impact, especially as experienced team members retire and newer developers step in. Instead of relying on formal handoffs or last-minute knowledge transfer, learning becomes part of the day-to-day work. A more modern development experience also makes it easier to bring in new talent and help existing teams work more effectively, which is becoming increasingly important as experienced developers retire. There are financial benefits as well, though they tend to follow rather than lead. As organizations adopt more flexible tooling and, in some cases, open-source solutions, they gain options. They are no longer as tightly bound to a single vendor or licensing model. Over time, that flexibility can translate into meaningful cost improvements. What Successful Organizations Do Differently Those outcomes don’t happen by accident. The organizations that get real value out of modernization tend to have leadership teams that approach it differently from the start. They don’t treat it as a tool decision or a one-time project. They treat it as an effort to improve how their environment operates, and they’re deliberate about how they go about it. That shows up in a few consistent ways: They get specific about the problem before looking for a solution. They take the time to determine why they’re modernizing before deciding how. Whether it’s speed, cost, talent, or competitiveness, that clarity shapes every decision that follows. A clearly defined objective keeps the effort grounded and helps teams prioritize what matters, measure progress, and avoid getting pulled in directions that don’t support the end goal.They take a hard look at how work flows today. Not how it’s documented or expected to work, but how it actually plays out in practice. That means mapping out the full path from development through deployment, including where work slows down, where approvals stack up, and where teams have created workarounds just to keep things moving. This step often surfaces issues that aren’t visible at a leadership level.They involve the people closest to the work. The most useful insights tend to come from the teams working in the process every day. Developers, operators, and support teams see where the friction is and what would make the biggest difference. Bringing those voices in early leads to better decisions and fewer surprises later.They establish clear ownership inside the organization. Modernization efforts move faster and more consistently when there’s a clear internal owner. Someone who understands the goal, can make decisions, and is accountable for keeping the work moving.They plan for adoption, not just implementation. Even when the technical work is straightforward, the transition isn’t. Teams need time to adjust to new workflows, learn new tools, and build confidence in the changes. Organizations that plan for that upfront tend to avoid the frustration that comes from trying to move too quickly.They start with a focused effort and build from there. Rather than trying to modernize everything at once, they begin with a smaller, well-defined scope. A pilot or targeted initiative creates a chance to test the approach, learn what works, and make adjustments before expanding more broadly. It also helps build internal support as people start to see tangible results. Making Modernization Work At its core, modernization isn’t about replacing one system with another. It’s about improving how the organization operates. Technology matters, but it only works when it’s built on a process that makes sense. Without that, modernization becomes another expensive layer on top of existing problems. When done well, modernization doesn’t just improve systems. It changes how teams work, how quickly the business can respond to what comes next, and turns a technical effort into a true business advantage.
TL;DR: The AI Definition of Done Your team has a Definition of Done for a product increment. It has none for the 20-plus AI-supported outputs that leave the team each week: status reports, stakeholder emails, release notes, and updates for the C-level. Each one carries your team’s name. “I know quality when I see it” is the standard most teams actually run by, and you cannot audit it, teach it to a new colleague, or defend it when a claim turns out to be wrong. The AI Definition of Done fixes that with one page per task class, agreed by the team, before the output ships. Your Increment Has a Standard; Does Your AI Output? A model turns the Jira board into a Friday status update, and the update tells an enterprise prospect that the security feature is in production. Unfortunately, it is not. The feature was descoped three months ago, but the old ticket title persisted because no one felt responsible. So the model reported the title instead of the reality. Nobody checked the claim against the release notes because nobody had agreed that someone should. The email was sent with the team’s name on the cover. A functioning agile team should be able to tell you what “done” means for a product increment. Few can tell you what “done” means for that status update. No agreed standard governs it, and it ships every week. The product increment passes through a standard that the team argued over and agreed on. The AI-assisted output passes through one person’s gut feeling at the moment they clicked send. One of those you can defend to a stakeholder, an auditor, or a new hire. The other you cannot. The AI Definition of Done closes that gap without adding a governance department, which is exactly why it survives in organizations where “AI governance” earns eye rolls. It takes a practice every agile practitioner already owns and points it at the work you have started handing to a model. It is not for everything: skip it for private brainstorming, throwaway prompts, or personal sensemaking, unless the output later informs a decision or leaves the team. The Four Questions Every AI Definition of Done Answers The Concept Verification Level Which claims get checked, by whom, against what source, and how? “Looks good” is not a method. A method names the claim, the checker, the source, and the test: every factual claim about product status gets checked against the release notes by the sender before sending, every time. Where teams get stuck: approval gets mistaken for review. Someone skims a draft, clicks send, and the team’s name now sits on a claim nobody verified. Provenance Disclosure What does the team declare about how the output was produced? Three labels cover practice: a) Human means no material AI contribution to the content, claims, or structure (a spellchecker does not count), b) AI-assisted means AI contributed to drafting, summarizing, or analysis, and a named human reviewed the output and decided, and c) AI-automated means AI produced and sent the output under predefined rules, without human review before release, audited at a set cadence. The line that matters runs through “reviewed”: clicking send on an unread draft is approval, never review. An output approved without reading is AI-automated, whatever the team tells itself. Data Hygiene What never enters a model on the way to this output? Name the exclusions concretely: personal data from team surveys, customer-identifiable information, anything your organization’s AI policy restricts. If the input rules in your A3 Handoff Canvas already cover this, point to them. Do not keep two versions of the same rule. Where teams get stuck: nobody wrote the exclusions down, so each person guesses, and the guesses differ. Sufficiency Tier and Environment Which model, plan, and data boundary are good enough for this task class, and why? A top-notch frontier model drafting calendar invitation may fail in this regard. The cheapest model, run locally on an old Mac mini, can write a board update but likely fails in the other. Capability is only half of it: a board update may need an enterprise plan with a no-training guarantee or an approved connector, even when a mid-tier model is plenty. If your team has a routing policy, point to the tier and the environment it mandates. If it does not yet, name the model and the plan, and explain in one sentence why both are enough. The AI Definition of Done Template Four questions, plus two operating controls, one page. Here is the template a team fills in per task class: DimensionYour Standard for This Task ClassTask classVerification level: What is checked, by whom, against what, howProvenance label: Human (Avoid) / Assist / Automate from the A3 Delegation Framework, and where the label appearsData hygiene: What never enters the modelSufficiency tier and environment: Wich model, plan, and data boundary, and why they are enoughSign-off: Who agreed, on what date, and the review dateStop rule: When the delegation is paused, downgraded, or returned to manual work The last two rows are operational, not definitional: Sign-off records who agreed and when, and the stop rule names the condition that pauses the delegation, because this standard should say not only when an output may ship but when the task class stops being eligible for AI at all. Without it, teams keep tuning the prompt or skill long after the delegation has proven unfit. A Worked Example: External Status Communication The status update failure that opened this article maps to one task class, status communication, leaving the company. Here is the team’s first AI Definition of Done for it: DimensionStandardTask classStatus communication leaving the companyVerification levelEvery claim about feature status is checked against the release notes by the sending manager, before sending, every timeProvenance labelAI-assisted; footer states “Drafted with AI, reviewed by [name]”; Assist is not permitted for this task classData hygieneNo customer names, no security-finding details, no internal financials enter the modelSufficiency tier and environmentMid-tier model on an enterprise plan with no model training; drafting from structured release data needs no frontier modelSign-offTeam agreed, dated; review after the next four status updatesStop ruleIf two updates in a review cycle need a factual correction after sending, the task class returns to manual drafting until the standard is revised The standard costs the sending manager about four minutes a week, set against an error that can put a flagship deal at risk. Write Your AI Definition of Done in 75 Minutes An AI Definition of Done that one person downloads and pastes into the wiki doesn’t change anything. The argument over the standard is where the standard takes hold. Run it as a workshop: Pick three task classes (10 minutes): Choose from work the team actually shipped in the last two weeks, never hypotheticals. The best candidates are outputs that leave the team.Draft in pairs (20 minutes): Each pair fills the template for one task class. Pairs work without comparing notes; divergence is the point.Argue the differences (25 minutes): Compare drafts. Where pairs disagree on verification level or provenance, the team has found an unspoken assumption. Resolve each disagreement with a decision, never with “both are fine.”Set the labels (10 minutes): Agree where provenance labels appear: email footers, document headers, report covers. Visible beats buried.Adopt and date (10 minutes): Sign off each AI Definition of Done with a review date, and add the adoption to your AI working agreement. Ownership stays with the team running the delegation. Compliance, security, or legal may constrain the standard, but they do not write it for the team. When someone says, “We do not need this for internal outputs,” ask what happened the last time an internal draft got forwarded outside the team. Every team has that story. The Record You Get for Free Each signed-off AI Definition of Done is a dated, versioned, one-page record. Stack them, and they answer the due diligence question enterprise buyers increasingly ask, “How do you control AI-generated output?” with documents instead of assurances. Nobody wrote a governance report. The records came out of normal work. That answer is already part of procurement and due diligence conversations. Article 4 of the EU AI Act has been applied since February 2, 2025, and requires providers and deployers to ensure a sufficient level of AI literacy among staff and others operating AI systems on their behalf. The EU Commission’s Q&A places supervision and enforcement under national market surveillance authorities, with the enforcement rules applying from early August 2026. The practical question underlying the regulation is simpler, and a prospect’s procurement team will ask it before any regulator does: can you show the standard that underlies the output you sent us? Three Ways It Fails The downloaded standard: A template adopted without the workshop. Nobody argued, so nobody owns it. An AI Definition of Done that nobody argued about is one nobody will follow. The universal standard: One AI Definition of Done for all work. Verification that aligns with external communication suffocates internal brainstorming, and the team abandons the practice within a month. One page per task class. Contrary to the classic Definition of Done, there is no one-size-fits-all in our use case. The static standard: Written once, reviewed never. Models change, people change, task classes change. The review date is part of the artifact, and your next delegation inspection enforces it. Conclusion: Pick One Output This Week Pick one AI-assisted output your team ships regularly. The Friday status update, the Sprint summary, or the stakeholder email. Walk it through the four questions out loud in your next Retrospective: what gets checked and by whom, how we label it, what never enters the model, and which tier is enough. You will likely find at least one question where the honest answer is “nobody decided that.” Write the one-page response for that task class, argue it, sign it, and date it. One standard, agreed by the team, is the difference between a team that uses AI and a team that a customer can trust with it. Which of your AI-assisted outputs has a standard behind it right now, and which one is merely a habit? Key Questions This Article Answers What Is an AI Definition of Done? An AI Definition of Done is a one-page, team-agreed standard that an AI-assisted output must meet before it leaves the team. Teams write one per task class, such as external status communication or data analysis summaries, never one per task. It answers four questions: what gets verified, how the output is labeled, what data never enters the model, and which model and environment are sufficient. It borrows the discipline of the Scrum Definition of Done and applies it to work on a model touched. What Is the Difference Between Approval and Review for AI Output? Review means a named human reads the AI-generated output and checks its claims against a source before it ships. Approval means someone clicked send. Clicking send on an unread draft is approval, not review, whatever the team calls it. An output approved without reading is effectively AI-automated, and it should carry that provenance label rather than the AI-assisted label, which implies a human verified it. How Do You Write an AI Definition of Done? Run a 75-minute team workshop, not a solo download. Pick three task classes from work shipped in the last two weeks, draft the standard in pairs, then compare and resolve every disagreement with a decision. Agree where provenance labels appear, set a stop rule that returns the task class to manual drafting when outputs repeatedly fail, sign off each standard with a review date, and add the adoption to your AI working agreement. The argument over the standard is what makes the team own it. How Do Agile Teams Prove They Govern AI Output? Each signed-off AI Definition of Done is a dated, one-page record. Together, a team’s standards answer the procurement and due diligence question “how do you control AI-generated output” with documents rather than assurances. The records are a byproduct of normal work, so no separate governance report is needed. This matters because buyers and regulators, including under the EU AI Act Article 4, increasingly require evidence of controlled AI adoption. What Are the Four Dimensions of an AI Definition of Done? Verification level (which claims get checked, by whom, against what source, and how), provenance disclosure (Human, AI-assisted, or AI-automated, and where the label appears), data hygiene (what never enters the model), and sufficiency tier and environment (which model, plan, and data boundary are good enough and why). Each dimension fits on one line of a one-page template, signed off with an adoption date and a stop rule that pauses the delegation when outputs repeatedly fail.
There is a pattern that repeats itself across engineering organizations regardless of team size, tech stack, or industry. A sprint ends. Features are shipped. The QA team is still writing automation for the previous sprint. The backlog of unautomated scenarios grows. Leadership asks what it would take to close the gap. The answer comes back: more engineers, more time, more tooling budget. Six months later, the gap is the same size. Sometimes larger. This is not a resource problem. It is an architectural problem. And until the architecture changes, the gap does not close. The Upstream Problem Nobody Measures When engineering teams analyze their automation coverage gaps, they almost always focus on execution test runs that are slow, maintenance is high, and flaky tests waste time. These are real problems. But they are downstream of a more fundamental issue that rarely gets measured: the time between a requirement being written and automation existing for it. In a traditional QA workflow, that gap looks like this: Requirement lands in JiraDeveloper builds the featureQA engineer reads the requirement, interprets it, designs test scenariosQA engineer writes test casesQA engineer scripts automation in Playwright or SeleniumQA engineer executes, debugs, maintains Steps 3 through 5 take days. Sometimes weeks. Every sprint adds to the backlog. Every requirement change breaks existing automation. The team runs hard and stays in the same place. The industry has responded to this by automating step 6, making execution faster, smarter, and more parallelized. But steps 3 through 5, requirement interpretation, test design, and scripting, remain almost entirely manual in most organizations. This is the upstream problem. And it is where the real automation opportunity sits in 2026. What Changes When You Start From Requirements The architecture shift that actually closes the coverage gap starts much earlier in the pipeline than most automation teams consider. Instead of "requirement arrives → developer builds → QA manually creates coverage," the new model is "requirement arrives → AI evaluates and enhances → AI generates test cases → AI generates scripts → AI executes → results with traceability returned." The human does not design coverage. The human does not script automation. The human reviews requirements, approves test cases when necessary, and focuses on exploratory testing and quality strategy, the work that actually requires human judgment. This is what requirement-driven autonomous testing means in practice. The requirement is the input. The executed test result is the output. AI owns everything in between. The 5 Stages of a Requirement-to-Result Pipeline Platforms like TestMax implement this model as a connected five-stage pipeline. Understanding each stage explains why the architecture works differently from traditional automation approaches. Stage 1: Requirement Ingestion The pipeline accepts requirements from wherever they live, Jira tickets, Azure DevOps work items, Word documents, PDFs, Excel files, or requirements authored directly in the platform. No reformatting required. The requirement enters the system as it exists. This matters because one of the friction points in traditional QA automation is the translation step, converting a Jira ticket into a format that test tooling can work with. When ingestion is native, that step disappears. Stage 2: Requirement Intelligence Before any test generation begins, every requirement is evaluated by AI across five quality dimensions: clarity, completeness, consistency, testability, and correctness. This stage is the most underestimated in the entire pipeline. Poor requirements produce poor tests always. A requirement that says "the login form should work correctly" is not testable. A requirement that specifies valid credentials, invalid passwords, empty field behavior, account lockout thresholds, and session persistence rules is. When AI catches ambiguity at the requirement stage, it costs nothing to fix. When that same ambiguity surfaces after automation has been built against it, it costs days. The requirement of the intelligence layer moves the defect detection upstream to where it is cheapest. Requirements that fail quality review are flagged with specific improvement suggestions. AI offers rewrites. Nothing ambiguous proceeds to test generation. Stage 3: AI Test Case Generation Once a requirement passes quality review, the platform generates structured test cases automatically. Not surface-level happy path scenarios, complete coverage across positive paths, negative paths, boundary conditions, and edge cases. For a single requirement, like users can reset their password via email verification, the generated coverage includes: Valid email address submitted – verification email receivedInvalid email format – appropriate error returnedEmail address not registered – system response without revealing account existenceVerification link clicked – password reset flow initiatedVerification link expired – appropriate error with re-send optionNew password does not meet policy requirements specific validation messagesSuccessful reset – session handling, redirect behaviour All of this is generated automatically from the requirement. No human designs the coverage strategy. Stage 4: Automation Generation Approved test cases are converted into executable Playwright scripts automatically. Production-ready code with appropriate waits, assertions, and selector strategies generated without a human writing a single line. This is the step that eliminates the scripting bottleneck. In traditional automation, scripting bandwidth is a hard ceiling on coverage growth. When the team can script 50 test cases per sprint, coverage grows at that rate regardless of how many requirements are produced. When scripts are generated automatically from approved test cases, that ceiling disappears. Coverage can grow at the rate requirements are produced, not the rate engineers can write code. Stage 5: Autonomous Execution and Evidence AI agents execute the generated test suite through Playwright MCP. They manage environment setup, handle retries, capture logs, screenshots, and video per test, and return a complete traceability matrix linking every result to its source requirement. The output is not a pass/fail count. It is a complete evidence package suitable for audit, governance, and release decision-making generated automatically from the requirements the team was already writing. Why This Architecture Closes the Coverage Gap The traditional automation model has a linear constraint: coverage grows proportionally to engineering effort. More requirements always mean more backlog because the human work required per requirement is roughly constant. The requirement-driven autonomous model removes the linear constraint. When AI handles test design, scripting, and execution per requirement, the engineering effort per requirement drops dramatically. Coverage can scale with the requirements themselves rather than with team headcount. There are three concrete consequences: Coverage lag is eliminated. When test generation takes minutes rather than days, new features can have automation in the same sprint they are built. The perpetual state of automation backlog, where coverage is always weeks behind the code it is supposed to validate, is a consequence of the manual model, not an inevitability. Maintenance burden shifts. In traditional automation, 60 to 80 percent of automation engineering effort goes to maintaining existing scripts. When AI generates scripts from requirements, the maintenance responsibility belongs to the generation layer. UI changes that would previously break dozens of handwritten selectors are addressed at the generation stage. Requirement quality improves as a side effect. When every requirement must pass an AI quality evaluation before entering the test pipeline, the incentive to write precise, testable requirements increases. Teams that implement requirement-driven testing typically report improvement in requirement quality within two to three sprints, not because they trained their product managers differently, but because the pipeline now provides immediate, specific feedback on every requirement. Integrating With Existing Workflows A practical concern with any architectural change is migration cost. The requirement-driven autonomous model does not require replacing existing infrastructure. Generated Playwright scripts integrate directly into existing CI/CD pipelines. Teams running Jira or Azure DevOps connect those systems natively requirements flow in without manual re-entry. For teams using ATF or other existing test frameworks, the autonomous testing layer runs alongside rather than replacing what already exists. The practical starting point is a single sprint. Take the new requirements entering your backlog this week. Run them through a requirement-driven platform. Compare the test coverage produced in time, in scenario depth, in maintenance overhead against what your team would have produced manually. The experiment answers the adoption question more convincingly than any benchmark. The Architectural Question for 2026 The relevant question for QA teams in 2026 is not whether to use AI in testing. Almost every serious testing platform has added AI capabilities in some form. The question is: where in the pipeline is AI actually doing meaningful work? At one end of the spectrum, AI heals broken selectors and suggests which tests to run. The human still reads requirements, designs coverage, writes scripts, and manages execution. AI makes individual tasks faster. At the other end, AI owns the pipeline from requirement evaluation through execution and evidence delivery. The human provides requirements and reviews results. AI does everything in between. The teams that figure out where they sit on that spectrum and decide consciously which model their coverage goals require are the ones that will stop having the same conversation about automation backlogs next quarter.
Teams often say they are building one app. A lot of the time, that is not true. I saw this while reviewing a telemedicine MVP. At first, the plan sounded simple enough: video visits, messaging, scheduling, and basic records. Then the version-one list kept growing: Patient appprovider dashboardAdmin panelMessagingVideoBillingEHR connectionDevice support later At that point, this was no longer one app. It was several systems being planned as one MVP. A patient-facing productA provider-facing productAn admin productA set of outside-service connections When a team treats all of that like one first release, things get messy before development even starts. The Moment It Stopped Being One App The problem was not the number of screens. The problem was the number of users, roles, and data rules hiding behind those screens. A patient needed intake, booking, reminders, and follow-up. A provider needed schedules, patient context, notes, and quick actions during the day. An admin needed visibility, support tools, and role controls. The outside-services side added video vendors, messaging vendors, EHR work, and, later, device data. That is not one product. That is a group of different systems with different jobs. Once that became obvious, the planning changed. Split the Product by User First Before estimating anything, it helps to split the product by who it is for. For this telemedicine project, the first useful split looked like this: 1. Patient Side This part handled: IntakeBookingRemindersFollow-up messagingJoining a visit The patient's side had to stay simple. It also had to be clear about what the patient could and could not see. 2. Provider Side This part handled: Schedule viewPatient detailsVisit notesQuick responsesRole-based access This was not just a different set of screens. It had different speed needs, different daily habits, and different data access rules. 3. Admin Side This part handled: Role setupSupport actionsVisibility into operationsReportingNon-clinical controls Admin work often looks small during planning. In real projects, it adds a lot of rules and a lot of testing. 4. Outside-Service Work This part handled: Video vendor setupMessaging vendor setupEHR-related workFuture device dataLogging and audit-related movement of data This is where many teams get surprised. Video, messaging, and EHR are not tiny add-ons. Each one brings its own work. Start With Access Rules Before the Feature List In multi-role products, one of the quickest ways to find hidden work is to define access rules early. Before locking the feature list, ask: Who can create this dataWho can read itWho can change itWho can delete itWho can export it For the telemedicine project, this made a big difference. A few features looked simple in the scope doc. Once the team asked who could view or change the related data, the work got much larger. A basic example: Admins can help fix booking problems. That sounds harmless. But then the real questions start: Can admins see messages?Can they see visit notes?Can they see call history?Can they open uploaded files? That one sentence can change a big part of the system. Access rules often show hidden work much faster than a feature list does. Treat Outside Services as Separate Work Another mistake teams make is treating outside services like small items on a checklist. On paper, it can look like this: VideoMessagingEHR later In practice, each one adds its own work: Vendor setupRequest and response formatsError handlingRetry rulesLoggingReplacement cost if the vendor needs to change later That is why these items should be planned separately. For the telemedicine case, once video, messaging, and EHR work were split out from the main product list, the first release became easier to define. Some items that seemed close to launch were clearly not ready for version one. Ship One Complete Path First Once the team stopped calling everything an MVP, the first release got smaller. The version-one path that stayed in looked like this: Patient intakeAppointment bookingSecure video through the chosen vendorFollow-up messagingBasic provider access controls That was enough to test whether the product solved a real problem for a clinic. What moved out of the first release: Deeper EHR workMore reportingDetailed billing flowsDevice supportBroader admin tooling Those things were not bad ideas. They just did not belong in the first build. 4 Simple Documents to Create Before Sprint Planning When a team starts to suspect that one MVP is several systems, four short documents can help a lot. 1. User-to-System Map List each part of the product and the main user for it. 2. Permission Matrix Write down who can create, view, change, delete, and export each type of data. 3. Outside-Service List Separate core product work from vendor work and data that moves in or out of the system. 4. First-Release Path Write the one end-to-end path that version one has to get right. These are short documents, but they make planning much better. Why This Matters Outside Healthcare, Too This lesson is not only for telemedicine. It applies to any multi-role product where the team is building for more than one type of user. That includes: Customer apps with admin panelsSaaS products with back-office toolsPlatforms with provider and client sidesProducts that depend on outside vendors from day one The moment a team has different users with different goals, the work stops being “just one app.” Final Point A lot of MVPs get too big because teams keep calling them one product long after that stops being true. The fix is not always better estimates. Sometimes the fix is much simpler: Split the product by user.Write down the access rules.Separate outside-service work.Ship one complete path first. That makes the first release easier to plan, easier to build, and easier to test.
In my 30 years of navigating the IT landscape, I’ve seen ‘Agile’ transform from a revolutionary mindset into what often feels like a series of manual project hurdles. In many large projects I’ve led, I’ve noticed we’ve traded innovation for a culture of ‘babysitting’ Jira boards and tracking Excel sheets. I wish to develop the Agentic Agile Office (AAO) not as another layer of automation, but as a fundamental shift in how I believe we must manage project velocity and governance. The Bottlenecks I’ve Encountered In my experience, traditional Enterprise Agile often buckles under its own weight. I’ve watched Technical Program Managers (TPMs) and Scrum Masters spend up to 60% of their time on administrative overhead. I’ve seen the "manual tax" of chasing status updates slow down the very speed Agile was designed to create. I believe it’s time to move past this. How I Define Autonomous AI Agents The AAO framework I’m proposing moves beyond simple chatbots. I am focusing on agentic AI — systems capable of reasoning, planning, and executing tasks autonomously. Within my framework, these agents don't just answer questions; they take action: The Backlog Agent: This will automatically analyze user feedback and technical debt to suggest prioritization scores for the Product Owner.The Dependency Agent: This agent scans multiple team boards in real-time. I want it to identify and flag architectural conflicts before they cause a sprint failure.The Governance Agent: I see this as the ultimate safeguard, ensuring all code commits meet compliance standards without a human auditor needing to manually check every pull request. Deep Dive: The Architecture of the AAO While defining these agents is the first step, I believe it is critical to understand the architectural engine that drives this office. To move beyond simple automation, I have structured the AAO as a three-tier system: 1. The Intelligence Layer: Reasoning Over Data In my three decades in the industry, the biggest issue hasn't been a lack of data, but the "data fog." I designed the AAO to use large action models (LAMs) that don't just read your tickets; they understand the intent behind them. Contextual memory: I want these agents to remember that a delay in a previous quarter was caused by a specific API bottleneck so they can predict similar risks today.Reasoning loops: Instead of a static trigger, I’ve structured these agents to use "Chain of Thought" processing to validate if a story is actually "Ready" based on historical standards. 2. The Workflow: A Day in the Life of an Agentic Sprint I’ve reimagined the standard sprint cycle to show exactly where I believe these agents provide the most value: Pre-planning: Before the team meets, I have the Backlog Agent scrub requirements. If a user story lacks an acceptance criterion, the agent flags it to the Product Owner immediately, saving us 30 minutes of "discovery" time during the meeting.In-sprint execution: I’ve implemented the Dependency Agent to act as a "digital scout." If a developer changes a schema that another team relies on, the agent detects the conflict in the pull request and notifies both Scrum Masters before the build even fails.The "always-on" retrospective: I believe retrospectives shouldn't just happen every two weeks. My Insight Agent tracks velocity trends daily. If I see a team's burndown stalling, the agent provides me with a root-cause analysis before I even ask. 3. My Strategy: Agentic Over Generative AI I want to be clear on a point of common confusion: Generative AI writes the email; agentic AI recognizes a project risk, decides an email is necessary, and drafts it for my review. In my framework, I am moving the human from being the operator of the tool to being the editor of the agent's actions. I’m shifting our workload from "doing the work" to "verifying the outcomes." Why I Believe This Redefines Our Roles This technical shift leads to a natural question: if agents are handling the logistics, what happens to the people? In my view, this shift doesn't diminish our roles; it elevates them. By offloading the "babysitting" of Jira boards to autonomous agents, I want to empower leadership to focus on: Complex problem solving: Negotiating high-level blockers that require a human touch.Mentorship: Spending more time coaching teams to improve their craft.Strategic alignment: Ensuring technical output truly maps to business value. My Vision for the Future To me, the Agentic Agile Office represents the transition from Agile-by-process to Agile-by-intelligence. I am confident that by integrating these agents, enterprises can finally achieve continuous delivery without the human burnout I’ve witnessed throughout my career. I no longer ask "How do we scale Agile?" I now ask: "How quickly can I help you integrate the agents that will do the scaling for you?"
Lead Engineer,
Technical University of Crete
Agile Coach,
Berlin Product People GmbH