DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Engineering AI Accountability: Control the Execution Path, Not the Model
  • What Full-Stack AI Engineering Means in Real Projects
  • Context Engineering: The Missing Piece in Agentic Systems
  • Why AI Hallucinations Are a Quality Engineering Problem

Trending

  • Decoding the “Black Box”: Evaluating Agent Tool Chains in Production
  • From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
  • How Go Maps Work: From Buckets to Swiss Tables
  • Building a Migration Readiness Engine for Analytics Workflows
  1. DZone
  2. Data Engineering
  3. AI/ML
  4. AI Has Solved the Code Bottleneck. Now Engineering Leaders Have a Measurement Problem.

AI Has Solved the Code Bottleneck. Now Engineering Leaders Have a Measurement Problem.

AI-assisted engineering is shifting the bottleneck from coding to verification. Learn why traditional metrics miss the hidden cost of AI-generated code.

By 
Igboanugo David Ugochukwu user avatar
Igboanugo David Ugochukwu
DZone Core CORE ·
Oct. 08, 26 · Analysis
Likes (0)
Comment
Save
Tweet
Share
155 Views

Join the DZone community and get the full member experience.

Join For Free

Picture the dashboard on a good day. Cycle time is green. Pull request counts are climbing. The throughput chart bends upward, exactly as the vendor deck promised. Now picture the person the dashboard cannot show you: a senior engineer spending her afternoon working out whether a flawless-looking pull request is actually correct.

That gap between what the chart says and what the reviewer feels is the story of AI-assisted engineering in 2026. Writing code is no longer the scarce resource. Knowing what the code does, what it cost, and whether anyone checked it is. Most organizations are still measuring the old scarcity.

What the Independent Evidence Shows

Start with the most rigorous public research on delivery performance. Google Cloud's 2025 DORA report drew on survey responses from nearly 5,000 technology professionals. It found that AI adoption now correlates positively with delivery throughput and product performance, but still correlates negatively with delivery stability. The authors' explanation is the technical heart of this whole topic: without robust control systems such as strong automated testing, mature version control practices, and fast feedback loops, a rise in change volume produces instability. The pipeline is receiving more input than its safeguards were sized for.

Developer sentiment moves the same way. Stack Overflow's 2025 developer survey, fielded from May 29 to June 23, 2025, with 49,009 responses according to ADTmag's coverage, found that 84% of developers use or plan to use AI tools, while trust in the accuracy of the output fell to 29% from 40% the year before. Adoption normally builds confidence. Here it did the opposite.

Then there is the perception problem. METR ran a randomized trial in 2025 with 16 experienced open source developers across 246 real tasks. Developers forecast a 24% speedup, and the measured result was that tasks took 19% longer, yet afterward they still believed AI had made them 20% faster. I want to be careful, because this result is routinely overstated. It covers one small group using early 2025 tools, METR itself now calls the results historical, and its 2026 follow-up was too compromised by selection effects to give a reliable estimate. The durable lesson is not that AI slows people down. It is that felt productivity and measured productivity can diverge, so self-report is a weak instrument for a measurement problem.

Corroboration From the Vendors, With Caveats

Two recent surveys come from companies that sell code verification tools, so read them as directional, not definitive.

Sonar's State of Code survey, released January 8, 2026, covered over 1,100 developers worldwide. Respondents reported that AI accounts for 42% of their committed code, with an expected 65% by 2027. Also, 96% said they do not fully trust AI output, yet only 48% said they always verify it before committing. The Register's coverage adds that 38% said reviewing AI code takes more effort than reviewing a colleague's, against 27% who said the opposite.

Qodo published its 2026 State of AI Code Quality Report on September 23. Censuswide surveyed 500 developers and 300 engineering leaders, all in the United States and all at organizations where AI already does meaningful work, between August 7 and August 14, 2026 (survey details). Developers and leaders, answering separate questionnaires, both named reviewing and validating AI-generated code as their top delivery constraint, at 26% each. The sample is not representative of all engineering organizations, and every figure is self-reported. Still, one result stands out. 90% of leaders said they can report AI's impact to executives or the board, while only 45% said they can trace AI activity to the code changes it produced. The same report found 89% of organizations had experienced an AI-related production incident.

Every one of these surveys measures perception, at a specific moment, with tools that keep changing. That limitation shapes what follows.

Why the Old Metrics Miss It

Three mechanisms explain why familiar dashboards mislead once AI writes a large share of the code.

The unit of measurement drifted from the unit of change. Tickets and story points record intent. They were a rough proxy for the amount of code that shipped, and that proxy held well enough when humans typed every line. When one engineer with an assistant can turn a multi-week refactor into an afternoon, the relationship between a ticket and the change it produced stops being stable.

Verification cost never appears on a clock. Qodo's report describes why. AI-authored changes tend to look finished, with clean naming and passing tests, but nothing in the diff shows which alternatives were considered or which assumptions carried over. Reviewers spend more effort reaching the same confidence, and 36% of developers in Qodo's sample said review takes the same time but demands greater cognitive effort. Cycle time cannot register effort that does not lengthen the calendar.

Feedback loops were sized for the old volume. This is DORA's finding restated. Review queues, test suites, and deployment safeguards were built for a certain rate of change. Raise the rate and the weakest safeguard becomes the constraint.

These mechanisms open three distinct measurement gaps. Spend cannot be tied to what got built. Activity rises without a matching rise in shipped value. And real effort, such as careful verification and cleanup of generated code, never gets a ticket.

Two Practitioner Perspectives

Flux, a Boston company building a code-first engineering intelligence platform, supplied commentary for this piece. Flux sells the kind of analysis its executives recommend, so weigh their views accordingly, but each speaks to a real part of the problem.

Ted Julian, Flux's founder and CEO, describes the pressure from the finance side:

"Every engineering leader we talk to has been doubling down on AI: more tooling and more code moving through the pipeline. And in nearly every enterprise conversation, the first questions from senior leadership are about spend. Can you show CapEx versus OpEx? Can you help substantiate an R&D credit?"

He argues that the organizations that cannot answer usually cannot see quality drift either, because both gaps come from "measuring activity instead of analyzing what the code actually shows." He has described Flux's code-based approach in more detail on The Lantern podcast.

Aaron Beals, Flux's CTO, speaks to the engineering side. "AI made these problems move faster than the old measurement systems can keep up with," he said, and the ticket describing the work and the codebase showing the result have drifted far enough apart to create serious blind spots. He ties two problems to one gap: Finance in the quarterly review "asking what the AI spend actually bought," and an engineer paged over a dependency change nobody reviewed. His proposed remedy, "Analyzing the codebase itself closes that gap for both questions without slowing anything down," is a vendor's claim, and I would test it in a pilot before believing it.

Whatever you conclude about any product, both executives point to the same requirement. The evidence has to come from what was built, not only from what was recorded about it.

A Measurement Framework

I find it useful to ask four questions about every AI-assisted change, and to keep each question's metrics in its own view.

What shipped? Merged changes and lead time. This is the layer most dashboards already have, and it should never appear alone.

Did it hold? Change failure rate, time to restore service, revert rate, and rework within thirty days of merge. DORA already defines the first two, so you can adopt them without inventing anything.

What did it cost to trust? Review rounds per change, comments per change, and time from first review to approval. These are proxies, and the honest label for them is "review effort," not "review quality."

What produced it, and what did it cost? Whether the change was AI-assisted, and how the spend maps to teams and repositories. This layer is the least mature in most organizations, and Qodo's 90% versus 45% gap suggests it is where confidence most outruns evidence.

Flux's briefing uses three labels for the resulting blind spots: unproven spend, velocity theater, and hidden work. They map cleanly onto this framework. Unproven spend is a gap in the fourth question. Velocity theater is the first question answered without the second. Hidden work is the third question left unmeasured.

Implementing It In 30 Days

  1. Baseline first. Pull the previous two quarters of lead time, change failure rate, time to restore, and revert rate for each repository before you change any process. An ROI claim without a before picture is an estimate.
  2. Mark provenance. Agree on a convention, such as a commit trailer or a pull request label, for AI-assisted changes. Expect it to be incomplete and partly self-reported at first, and cross-check it against any usage data your tools expose.
  3. Split the dashboards. Put throughput on one view and stability and rework on another. Do not average them into a single velocity score.
  4. Compare like with like. Within the same team and repository, compare AI-assisted and unassisted changes on rework and escaped defects. Comparing across teams mostly measures the teams.
  5. Map spend to work. Attribute licenses and usage to teams and repositories, and hand Finance that mapping. Questions such as CapEx versus OpEx treatment and R&D credit eligibility belong to your finance and tax advisors, and code-level evidence is an input to their judgment, not a substitute for it.

Where This Approach Can Fail

Any metric becomes a target once people know it is watched, so pair each measure with its counterweight, and review the set regularly. Provenance tagging depends on honest, consistent use and on what your tools can report. The DORA findings and the surveys are correlational, and cohort comparisons inside one company are not randomized experiments. Finally, all of this evidence describes tooling from 2025 and early 2026, and agentic workflows are changing quickly. Treat the framework as something to recalibrate, not to install once.

Disclosure

Sonar and Qodo sell verification and code review products and produced the surveys cited above. Flux supplied the executive commentary and sells code-first engineering analytics. The independent sources, DORA, Stack Overflow, and METR, point in the same direction, but each has its own limits, noted above.

AI removed the bottleneck at the keyboard and moved it to trust. Trust is built from evidence about what the code does, and a ticket cannot supply that.

Sources

  1. Google Cloud, Announcing the 2025 DORA Report:
    https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
  2. Stack Overflow, 2025 Developer Survey for Leaders:
    https://stackoverflow.co/internal/resources/2025-stack-overflow-developer-survey-for-leaders/ai-adoption/
    ADTmag coverage with fieldwork dates:
    https://adtmag.com/blogs/watersworks/2026/01/stack-overflow-survey.aspx
  3. METR, Early 2025 developer productivity study:
    https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
    METR, February 2026 update:
    https://metr.org/blog/2026-02-24-uplift-update/
  4. Sonar, State of Code Developer Survey report (January 8, 2026):
    https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding/
  5. The Register, Devs doubt AI-written code, but don't always check it:
    https://www.theregister.com/software/2026/01/09/devs-doubt-ai-written-code-but-dont-always-check-it/4932910
  6. Qodo, 2026 State of AI Code Quality Report (September 23, 2026):
    https://www.qodo.ai/blog/state-of-ai-code-quality-report-2026/
    GlobeNewswire release with survey dates:
    https://www.globenewswire.com/news-release/2026/09/23/3367496/0/en/qodo-s-2026-state-of-ai-code-quality-report-reveals-growing-verification-challenge-as-agentic-development-scales.html
  7. Flux, About page:
    https://www.askflux.ai/about/
  8. MGMT Boston, Ted Julian on The Lantern:
    https://mgmtboston.com/lantern/ted-julian-flux
AI Engineering

Opinions expressed by DZone contributors are their own.

Related

  • Engineering AI Accountability: Control the Execution Path, Not the Model
  • What Full-Stack AI Engineering Means in Real Projects
  • Context Engineering: The Missing Piece in Agentic Systems
  • Why AI Hallucinations Are a Quality Engineering Problem

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook