• About us
  • Services
  • Careers
  • Blog
  • Home
  • -
    Blog
  • -
    What Engineering Leaders Should Actually Measure When AI Writes Most of the Code
Article Content
  • Chapter 1.Key Takeaways
  • Chapter 2.Introduction
  • Chapter 3.The evidence: what 2026's research actually says
  • Chapter 4.Why this is happening: the metrics literature solved a different problem
  • Chapter 5.What most engineering organisations get wrong
  • Chapter 6.Business and technology implications for BFSI
  • Chapter 7.A decision framework: The Custody Chain
  • Chapter 8.Counterargument and nuance
  • Chapter 9.The sourceCode perspective
  • Chapter 10.Conclusion
  • Chapter 11.FAQ
  • Chapter 12.Reference List

What Engineering Leaders Should Actually Measure When AI Writes Most of the Code

sourceCode | BFSI Technology Insight | 7 October 2026

Key Takeaways

  • 2026's most rigorous developer-productivity research - DORA's 2025 report, GitClear's code-quality analysis of 623 million changes, and MIT-adjacent lab METR's randomised trial - converges on one uncomfortable finding: AI adoption is nearly universal (90% of developers, per DORA) but its effect on delivery stability and code maintainability is negative, even as perceived productivity rises.

  • None of these three research programmes - the three most-cited bodies of AI-coding evidence in 2026 - were designed to answer a BFSI-specific question: which lines of a production system were AI-generated, under which model version, and reviewed by whom. That gap is real, but it is a gap in the tooling, not yet in explicit regulatory text.

  • APRA's 30 April 2026 letter to industry and ASIC's 8 May 2026 cyber resilience letter both push, in principles-based language, toward exactly this kind of lifecycle accountability and auditability - without yet naming "AI-generated code" as a specific disclosure category. Engineering leaders who wait for that explicit line item will be building the evidence base retroactively, under supervisory pressure, rather than by design.

  • Generic AI-coding metrics (acceptance rate, PRs merged, lines suggested) measure adoption, not exposure. They cannot answer "what changed in the credit-decisioning engine last quarter, and who is accountable for it" - the question a prudential regulator, an internal auditor, or a post-incident review will actually ask.

  • We introduce a four-part decision framework - the Custody Chain - for engineering leaders who need to close that gap before it becomes an examination finding rather than a design choice.

When Engineering Leader Should Actually Measure 2 1.png

Introduction

By the middle of 2026, the debate over whether AI coding tools make developers faster has produced a genuinely strange result: the two best-designed studies of the year disagree with each other, and both are probably right, because they're measuring different things. DORA's 2025 State of AI-assisted Software Development report - built on survey responses from nearly 5,000 technology professionals plus qualitative interviews - found that AI adoption has flipped from a negative to a positive correlation with software delivery throughput, while its correlation with delivery stability remains negative (DORA, 2025). METR's randomised controlled trial of 16 experienced open-source developers across 246 real tasks found the opposite at the individual level: developers using AI tools took 19% longer to complete tasks than developers who didn't, even though those same developers - before, during, and after the experiment - believed AI had made them roughly 20% faster (METR, 2025).

Put those two findings next to GitClear's analysis of 623 million code changes between 2023 and 2026 - which found code duplication up 81%, refactoring activity down from 13% to 3.8% of changed lines, and copy-pasted code climbing from 9.4% to 15.7% of all changes (GitClear, 2026) - and a pattern emerges that has nothing to do with whether AI is "good" or "bad" for developers. It has to do with what gets measured. Organisations chasing velocity signals (PRs merged, suggestions accepted, story points closed) are seeing exactly the productivity story they're measuring for. Nobody in that measurement stack is tracking what actually accumulates in the codebase, who is accountable for it, or - critically for a bank, insurer or super fund - which model, which version, and which reviewer stood behind each material change.

This is the gap this piece addresses. Not "is AI coding good," but: for an engineering organisation operating inside APRA's and ASIC's supervisory perimeter, what should actually be on the dashboard - and how far is the industry's current metrics literature from answering that question?

The evidence: what 2026's research actually says

Three bodies of work anchor this discussion, and it is worth being precise about what each one actually measured, because secondary coverage of all three has been loose in ways that matter.

Three Studies. Three Layers_ 1.png

DORA 2025 surveyed developers and organisations, not code repositories. Its headline figures - 90% of respondents report using AI at work (a 14-point rise year-on-year), over 80% report a productivity increase, and around 30% report little or no trust in AI-generated code - describe sentiment and self-reported outcomes, not measured throughput (DORA, 2025; Google Cloud, 2025). The report's most consequential finding is structural, not statistical: AI adoption is now positively associated with software delivery throughput, but continues to show a negative relationship with delivery stability. DORA's interpretation, introduced through a new "AI Capabilities Model" of seven organisational capabilities, is that AI functions as an amplifier - it magnifies whatever discipline (or lack of it) already exists in an organisation's testing, version control and platform engineering practices, rather than substituting for that discipline.

METR's trial is the more rigorous of the two on causal inference, precisely because it's smaller and narrower. Sixteen experienced developers, working in codebases they already knew well, were randomly allowed or disallowed AI tool use across 246 real GitHub issues. The measured result was a 19% increase in completion time when AI was permitted - the opposite of both the developers' pre-registered expectation (a 24% speedup) and their post-hoc belief (a 20% speedup) (METR, 2025). This is the sharpest available evidence that self-reported productivity gains - the kind that populate most vendor case studies and, frankly, most internal engineering-leadership dashboards - should not be taken as ground truth without an independent check.

GitClear's codebase analysis, drawn from 623 million analysed changes across the 2023-2026 period, is the closest thing available to a direct measurement of what AI assistance is doing to the code itself rather than to developer sentiment. Its clearest signal is a decline in refactoring: the share of changed lines classified as "moved" (indicative of code being restructured and reused rather than rewritten) fell from 13% in 2023 to 3.8% year-to-date in 2026, while block-level code duplication rose 81% over the same window (GitClear, 2026). A caveat is necessary here: GitClear's earlier 2025 report, covering 211 million lines through 2020-2024, was widely summarised in secondary coverage as showing "4x more code cloning." The underlying figure GitClear itself reports is a rise in cloned lines from 8.3% to 12.3% of changes - a real and material increase, but roughly a 48% relative increase, not 4x (GitClear, 2025). We flag this because it is a documented case of a primary statistic being compressed into a much more dramatic-sounding secondary claim, and it is worth checking the source directly before repeating it.

Read together, these three bodies of evidence support a specific, narrower claim than the one often made in 2026 conference talks: AI coding tools are increasing throughput at the organisational level (per DORA's survey data) while simultaneously increasing code churn, duplication and rework risk at the codebase level (per GitClear) and, in at least one controlled setting, decreasing task-level efficiency for experienced developers working in familiar code (per METR). None of the three studies were designed to, or do, measure code provenance - which model produced which lines, at what confidence, reviewed by whom, under what test coverage. That is a distinct and currently under-instrumented question.

Why this is happening: the metrics literature solved a different problem

The 2023-2026 shift in developer-productivity thinking - from DORA's four keys (deployment frequency, lead time, change failure rate, time to restore) toward DevEx - was a genuine and well-evidenced correction, not a rebrand. Noda, Storey and colleagues' 2023 ACM Queue paper "DevEx: What Actually Drives Productivity" argued, based on empirical work with developers across multiple organisations, that output-based and time-based productivity measures fail to capture the actual drivers of developer effectiveness, and proposed three dimensions instead: feedback loops (the speed and quality of response to a developer's actions - code review latency, build times, test feedback), cognitive load (the mental effort required to complete a task, driven by unnecessary complexity, poor documentation, and context-switching), and flow state (the ability to sustain focused, uninterrupted work) (Noda, Storey et al., 2023). This is a legitimate and useful correction to a real problem: DORA's four keys, applied naively, can be gamed by teams that ship small, frequent, low-risk changes while avoiding anything architecturally significant.

But DevEx was built to answer "why do developers feel slow, frustrated, or blocked" - a workplace-experience question. It was never built to answer "can this organisation demonstrate, to a prudential regulator or an internal audit function, who is accountable for a specific change in a production system." Those are different questions with different stakeholders. A DevEx survey score can improve at the same time as code provenance risk rises, because nothing in the DevEx framework, or in the AI-coding-ROI benchmarks that followed it through 2025 and 2026, asks the codebase to answer for itself. The 2026 developer-productivity literature - DORA's capabilities model, DevEx, and the wave of AI-ROI benchmark studies - has almost entirely inherited a framing built for a US technology-sector audience where "faster, happier developers" is close to the whole story. BFSI engineering leaders are operating a layer beneath that: velocity and developer sentiment matter, but they are necessary, not sufficient, conditions for what a regulator, an auditor or a court will eventually ask for.

What most engineering organisations get wrong

The most common mistake we see is not ignoring governance - it's assuming that a mature DORA/DevEx measurement practice is governance-ready by default. It isn't, for three specific reasons.

Where Attribution Gets Lost 1.png

First, acceptance-rate and adoption metrics (how many AI suggestions were accepted, what percentage of a PR's lines were AI-authored at commit time) are typically captured at the IDE or Copilot-style tool layer, and that data frequently does not survive into the merged codebase's permanent record. Once a PR is squashed and merged, the attribution metadata that existed at authorship time is very often gone. An organisation can have excellent AI-adoption telemetry and zero ability to answer, eighteen months later, which lines in a specific production incident trace back to an AI suggestion.

Second, "review happened" is treated as equivalent to "review was proportionate to risk." A human reviewer approving a PR is not the same signal as a human reviewer who understood, and specifically tested, the AI-generated logic inside a change to a credit-decisioning or AML-screening system. DORA's own 2025 findings support this concern directly: the negative correlation between AI adoption and delivery stability persists precisely in organisations without strong automated testing and mature version control - the review step is present, but the surrounding discipline that would make that review meaningful is not (DORA, 2025).

Third, engineering leaders often treat "model version" as an implementation detail rather than a governed input. APRA's 30 April 2026 letter to industry is explicit that boards need "ownership and accountability across the AI lifecycle, from design and development through to deployment, monitoring and decommissioning," and separately warns against "overreliance on vendor presentations and summaries without sufficient examination of key AI risks" (APRA, 2026). A coding assistant that silently upgrades its underlying model - a routine SaaS event for most AI coding tools - is exactly the kind of lifecycle event APRA's letter is describing, and very few engineering organisations currently log which model version authored which committed change.

Business and technology implications for BFSI

The regulatory direction of travel in Australia in 2026 has been unusually concrete for what is still, formally, a principles-based regime. APRA's April letter followed direct 2025 engagement with larger ADIs, insurers and superannuation trustees and reported that assurance practices are lagging behind the pace of AI implementation across the industry - not a hypothetical risk, an observed one (APRA, 2026). It calls for "continuous validation or monitoring... to detect issues such as model drift, bias, failure modes, or control breakdowns in a timely manner," and for "contractual and governance arrangements which provide sufficient transparency, auditability and assurance over AI services," extending visibility to "material, third-party and fourth-party dependencies" (APRA, 2026). ASIC's parallel letter, issued 8 May 2026, treats AI-accelerated cyber risk as a core licensing obligation requiring board-level attention, not a technical footnote (ASIC, 2026). Separately, CPS 230 - APRA's operational risk standard, in force since 1 July 2025, with material service provider contract terms due to be fully compliant by 1 July 2026 - already requires a Material Service Provider register and contractual assurance terms for any provider supporting a critical operation (APRA, 2025). An AI coding platform embedded in the delivery pipeline for a critical system is a plausible candidate for that register, whether or not an organisation has classified it that way yet.

None of this amounts to a regulator explicitly asking, today, for a line-by-line AI-authorship ledger inside production code. It is important to be precise about that, because overstating current regulatory specificity is its own credibility risk. What the evidence does support is a narrower, still consequential claim: the auditability, lifecycle-ownership and vendor-transparency obligations already in force, or already signalled as supervisory priorities, will be very difficult to satisfy retroactively if an engineering organisation cannot show, for a system under review, what proportion of it was AI-generated, by which model and version, and under what review standard. The cost of building that evidence trail after an incident or an examination request is materially higher than the cost of instrumenting it at commit time.

A decision framework: The Custody Chain

We propose a four-part framework - the Custody Chain - as a practical complement to (not a replacement for) DORA and DevEx metrics. It is designed to answer the specific question those frameworks don't: for any material system, can you account for what AI touched, and who stands behind it.

The Custody Chain_1 1.png

1. Origin. For any committed change of material significance, is the generating model and version identifiable at merge time, not just at IDE-suggestion time? This requires attribution metadata to survive squash-and-merge workflows, which most default Git tooling does not preserve without deliberate configuration.

2. Proportion. What share of a file, service or system is AI-originated, tracked over time rather than as a point-in-time snapshot? A system that was 5% AI-authored a year ago and is now 60% AI-authored has undergone a material change in its risk profile even if every individual PR was reviewed and merged normally.

3. Review weight. Was the human review proportionate to the system's criticality - meaning, for a change to a credit-decisioning, AML, payments or capital-calculation system, was there evidence of independent testing and domain-specific scrutiny, not just approval? This is the point where DORA's stability findings and APRA's assurance expectations converge: a review step that exists without corresponding test rigor is a documented risk factor, not a control.

4. Standing assurance. Is there ongoing, not point-in-time, monitoring for the specific failure modes AI-generated code introduces - duplication and unreviewed rework accumulating in a codebase (as GitClear's data shows is happening broadly), or a vendor's silent model upgrade changing the behaviour of a tool already embedded in the delivery pipeline (the scenario APRA's letter treats as a lifecycle-ownership gap)?

The Custody Chain is not a replacement for velocity or DevEx metrics - an organisation still needs to know if its developers are fast, unblocked and not burning out. It is the layer that sits alongside those metrics for any system where "we shipped it quickly and it was reviewed" is not, on its own, an answer a regulator, auditor or board risk committee will accept.

Counterargument and nuance

The strongest objection to this piece's framing is a fair one: no Australian regulator has, as of this writing, published a rule requiring line-level AI-code attribution, and it would overstate the evidence to claim otherwise. APRA's and ASIC's 2026 letters are explicitly framed as applying existing principles-based obligations to AI, not introducing new prescriptive requirements (APRA, 2026; ASIC, 2026). An engineering leader could reasonably argue that building granular provenance infrastructure ahead of an explicit requirement risks over-engineering a control that regulators may never specify in the form this piece anticipates.

That objection has merit, and it is worth taking seriously rather than dismissing. Our response is not that the requirement exists today, but that the direction is already legible from APRA's own language - "ownership and accountability across the AI lifecycle," "auditability... over AI services," continuous monitoring for "control breakdowns" - and that principles-based regimes have a consistent historical pattern in Australian financial services: the specific evidentiary expectations tend to arrive first through supervisory findings and enforcement action following an incident, not through a prescriptive standard published in advance. An organisation that has already instrumented origin, proportion, review weight and standing assurance is answering a question before it's asked in exactly the form regulators have signalled it will eventually be asked. An organisation that hasn't is building the same evidence base later, under time pressure, likely during an incident review - a materially worse position, even though the underlying rule never changed.

The sourceCode perspective

In our engagements with engineering teams across banking, insurance and super, the AI-coding conversation almost always starts and stops at velocity: acceptance rates, PRs per sprint, developer sentiment scores. Those are legitimate, useful numbers, and we don't think teams should stop tracking them. But we've also seen, directly, how quickly the conversation changes once an incident post-mortem or an internal audit asks a version of "show us what changed in this system and who's accountable for it," and the honest answer is that nobody can reconstruct it cleanly because the AI-attribution data lived only in an IDE session that no longer exists. That's not a governance failure of malice - it's an instrumentation gap, and it's fixable at relatively low cost if it's designed in before the system in question becomes the subject of a review.

The organisations we've seen handle this well didn't wait for a mandate; they treated origin and review-weight tracking as a normal extension of the code review tooling they already had, not a new bureaucratic layer bolted on top of engineering. Talk to us here!

Conclusion

The 2026 developer-productivity literature has genuinely moved past a narrow focus on DORA's four keys, and DevEx's three-dimension framework - feedback loops, cognitive load, flow state - is a well-evidenced and useful correction for teams trying to understand why AI adoption doesn't automatically translate into better delivery outcomes. But that correction was built to solve a developer-experience problem, not a supervisory-accountability one, and the evidence available in 2026 - from DORA, GitClear and METR alike - describes velocity and code-quality trends without touching provenance at all.


For an engineering organisation operating inside APRA's and ASIC's perimeter, the gap between those two questions is not hypothetical, and the regulatory language already published in 2026 makes the direction of travel reasonably clear, even without a prescriptive rule yet in place. The Custody Chain - origin, proportion, review weight, standing assurance - is offered as a starting structure for closing that gap deliberately, rather than reconstructing it under pressure later.

FAQ

Is AI coding actually making developers faster, based on the 2026 evidence? It depends on what's measured. DORA's 2025 survey data shows a positive correlation between AI adoption and organisational delivery throughput. METR's controlled trial of experienced developers on familiar codebases found a 19% increase in task completion time when AI use was permitted. Both can be true simultaneously because they measure different things at different levels - self-reported organisational throughput versus measured individual task time.

Is the "4x more code cloning" GitClear statistic accurate? The figure commonly cited in secondary coverage overstates GitClear's own reported number. GitClear's 2025 report states cloned code rose from 8.3% to 12.3% of changed lines - a real and material increase, but closer to 48% relative growth than 4x. Its more recent 2026 report, covering a larger dataset, reports block-level duplication up 81% over 2023-2026, using a different measurement method. We'd recommend citing GitClear's own reported percentages directly rather than repeating the "4x" framing.

Has APRA or ASIC specifically required AI code provenance tracking? Not explicitly, as of this writing. Both regulators' 2026 letters apply existing principles-based obligations - lifecycle accountability, auditability, third-party transparency, continuous monitoring - to AI generally, without naming source-code attribution as a specific requirement. The case in this piece is that those existing obligations make code-level provenance a reasonably foreseeable evidentiary expectation, not that it is already codified.

Does this apply only to systems that are fully AI-generated? No. The relevant risk threshold is materiality of the system, not the proportion of AI-generated code within it. A credit-decisioning engine that is 15% AI-authored still needs proportionate review and attribution for that 15%, particularly where GitClear's data suggests AI-associated changes carry a higher duplication and lower-refactoring profile than human-authored changes.

Where should an engineering leader start if none of this is currently tracked? Start with Origin - configuring commit and PR tooling to preserve model/version attribution through to merge, for systems above a defined criticality threshold. That single change unlocks most of the rest of the Custody Chain retrospectively for future changes, even if historical attribution can't be reconstructed.

Reference List

APRA (2025) Prudential Standard CPS 230 Operational Risk Management. Available at: https://www.apra.gov.au/operational-risk-management (Accessed: 7 October 2026).

APRA (2026) Letter to Industry on Artificial Intelligence (AI), 30 April. Available at: https://www.apra.gov.au/news-and-publications/apra-letter-industry-artificial-intelligence-ai (Accessed: 7 October 2026).

ASIC (2026) 26-092MR ASIC Calls for Urgent Cyber Uplift as AI Accelerates Cyber Threats, 8 May. Available at: https://www.asic.gov.au/about-asic/news-centre/find-a-media-release/2026-releases/26-092mr-asic-calls-for-urgent-cyber-uplift-as-ai-accelerates-cyber-threats (Accessed: 7 October 2026).

DORA (2025) State of AI-assisted Software Development 2025. Available at: https://dora.dev/dora-report-2025/ (Accessed: 7 October 2026).

GitClear (2025) AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones. Available at: https://www.gitclear.com/ai_assistant_code_quality_2025_research (Accessed: 7 October 2026).

GitClear (2026) The Maintainability Gap: 2026 AI Code Quality Research. Available at: https://www.gitclear.com/the_ai_code_quality_maintainability_gap (Accessed: 7 October 2026).

Google Cloud (2025) Announcing the 2025 DORA Report. Available at: https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report (Accessed: 7 October 2026).

Google (2025) How Are Developers Using AI? Inside Google's 2025 DORA Report. Available at: https://blog.google/innovation-and-ai/technology/developers-tools/dora-report-2025/ (Accessed: 7 October 2026).

METR (2025) Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Available at: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ (Accessed: 7 October 2026).

Noda, A., Storey, M.A. et al. (2023) 'DevEx: What Actually Drives Productivity: The Developer-Centric Approach to Measuring and Improving Productivity', ACM Queue, 21(2). Available at: https://cacm.acm.org/practice/devex-what-actually-drives-productivity/ (Accessed: 7 October 2026).

Related articles

25/09/2026

A CFO's Guide to the Real Cost of an AI Pilot That Never Scales

24/09/2026

Reinsurance Is Quietly Becoming the Testing Ground for Agentic AI in Insurance

01/10/2026

What the First Results From Hong Kong's Bank AI Sandbox Actually Show

07/10/2026

What Engineering Leaders Should Actually Measure When AI Writes Most of the Code

01/10/2026

Why the UAE's New Central Bank Law Also Regulates Your Technology Partner

28/09/2026

What "Explainable AI" Actually Needs to Mean for an Underwriting Decision

30/09/2026

Why the Next Wave of BFSI Technology RFPs Will Ask Different Questions

23/09/2026

The Hidden Cost of Shadow AI in Financial Services Back Offices

29/09/2026

The Case for Treating Data Residency as a Product Decision, Not a Legal One

06/10/2026

How CFOs Are Actually Allocating the 2026 AI Budget (And What It Means for Vendor Selection)

Navigating the Future of Software

linkedin
About usResources
SolutionssBrainChatbotVoicebotVoice RecognitionFace Recognition
Blog and InsightsAI & Blockchain Trends Industry Case Studies Thought Leadership Articles Success Stories & Client Spotlights 
Legal Privacy Policy Terms of Service 
linkedin

Australia - Malaysia - Vietnam

Copyright © 2026 source[code].

Australia - Malaysia - Vietnam