• About us
  • Services
  • Careers
  • Blog
  • Home
  • -
    Blog
  • -
    The Hidden Cost of Shadow AI in Financial Services Back Offices
Article Content
  • Chapter 1.Key Takeaways
  • Chapter 2.Introduction
  • Chapter 3.The evidence: shadow AI is not an edge case, it is the default
  • Chapter 4.The buyer's decision problem
  • Chapter 5.What most institutions get wrong
  • Chapter 6.The Desire Path Protocol
  • Chapter 7.Business and technology implications
  • Chapter 8.A fair counterargument: bans have worked, at least for a while
  • Chapter 9.What leaders should do next
  • Chapter 10.The sourceCode’s perspective
  • Chapter 11.Conclusion
  • Chapter 12.Frequently Asked Questions
  • Chapter 13.Reference List

The Hidden Cost of Shadow AI in Financial Services Back Offices

Key Takeaways

  • Across surveyed organisations, MIT's Project NANDA found regular personal AI tool use among staff at over 90% of companies, against official LLM subscriptions at only 40% - a 50-point gap between real usage and governed usage.

  • In UK financial services specifically, 37% of employees say they often use public AI tools like ChatGPT or Copilot for work, and 55% have never received formal training on how to do so safely (Smarsh, 2025).

  • Only 12% of financial firms using AI have adopted a formal AI risk management framework, and 92% have no policy governing third-party AI use at all (ACA Group and NSCP, 2024).

  • Banning consumer AI tools has a track record in banking - JPMorgan, Citigroup, Goldman Sachs and others did it in 2023 - but most quietly reversed course once they had a sanctioned alternative to offer instead.

  • The fix is not a better policy memo. It is a measurable deployment speed: how fast a discovered shadow AI use case becomes a governed, sanctioned one. sourceCode calls this diagnostic and response model the Desire Path Protocol.

shadow-ai-financial-services-governance

Introduction

Most conversations about generative AI risk in financial services start from the premise that the organisation is deciding whether to adopt AI. That premise is out of date. The decision has already been made - by underwriters, claims handlers, credit analysts and relationship managers who opened a browser tab, pasted in a policy document or a customer email, and got an answer faster than any sanctioned system could give them.

MIT Project NANDA (2025): companies with an official LLM subscription vs. employees already using personal AI tools for work

This is not hypothetical. Research from MIT's Project NANDA, published in August 2025, found that while only 40% of the companies it studied had purchased an official large language model subscription, staff at more than 90% of those same companies reported using personal AI tools regularly for work - often "multiple times a day, every day," while the official pilot sat stalled. That gap is the subject of this piece.

The prevailing narrative about enterprise AI failure - that most sanctioned pilots never reach production - is true and well documented. What it obscures is the mirror image: when the sanctioned path is too slow, work does not stop. It moves to whatever tool is already open on the desktop. For a bank, insurer or fintech handling customer financial data, regulated advice and audit-bound decisions, that is a governance blind spot most institutions cannot currently see, let alone size.

This article sets out the evidence, the decision problem it creates for CTOs and risk leaders, a fair look at the alternative - banning the tools outright - and a practical framework, the Desire Path Protocol, for closing the gap deliberately rather than by decree.

The evidence: shadow AI is not an edge case, it is the default

Three independent evidence sources, each with disclosed methodology, converge on one conclusion: unsanctioned AI use in professional settings is now majority behaviour, not a fringe one.

Four gaps that open when back-office work runs through unsanctioned consumer AI tools

MIT Project NANDA (2025) surveyed 153 senior leaders, ran 52 structured organisational interviews and analysed more than 300 publicly disclosed AI initiatives. Its widely cited finding - that roughly 95% of enterprise generative AI pilots fail to show a measurable P&L return - describes the sanctioned side of the ledger. The less-discussed finding describes the unsanctioned side: a "shadow AI economy" in which employees use personal ChatGPT and Claude accounts to automate real portions of their jobs, often without IT's knowledge, because those tools already do what the official rollout was meant to do - only faster, and without a change request.

KPMG International and The University of Melbourne's 2025 global trust study, based on 48,340 respondents across 47 countries - the largest study of its kind currently available - found that nearly half of working respondents admit to using AI in ways that contravene their organisation's policies, and that over half actively avoid disclosing their AI use or present AI-assisted work as entirely their own. Three in five say they have witnessed a colleague doing the same. The study also found governance and training lagging adoption almost everywhere surveyed - this is not primarily a discipline problem, it is a readiness gap.

Sector-specific data narrows this further. A Smarsh-commissioned, OnePoll-fielded survey of 2,000 UK financial services and insurance employees (April 2025) found 37% often use public AI tools such as ChatGPT or Copilot for work, 38% do not know whether their firm captures AI tool outputs, and 55% have never received official training on AI tool use. Separately, ACA Group and the National Society of Compliance Professionals surveyed over 200 compliance leaders at financial firms in mid-2024 and found 92% had no policy governing third-party AI tool use, and just 12% of firms actively using AI had adopted a formal AI risk management framework.

Read together, these are not four disconnected data points. They describe one condition: adoption has outrun governance, employees are not disclosing it, and most financial institutions have no reliable way to see how much back-office work already runs through infrastructure they do not control, log or audit.

The buyer's decision problem

For a CTO or Chief Risk Officer, this is a genuinely hard problem, not a lazy one. Three constraints collide:

1. You cannot govern what you cannot see. Standard controls - a workstation policy, an acceptable-use clause, an annual training module - depend on employees self-reporting behaviour that the KPMG data shows they are, on average, actively concealing.

2. The productivity case is real, which is exactly the problem. The US Federal Reserve Bank of St Louis (Bick, Blandin and Deming, 2025), using the first nationally representative US survey of generative AI adoption, found 28% of workers already using generative AI at work and self-reported average time savings of 5.4% of hours among active users. Staff are not routing around governance out of recklessness - they are responding rationally to a tool that works.

3. Sanctioned deployment inside regulated financial services is, structurally, slow - vendor due diligence, data residency checks, model risk sign-off, third-party risk assessment. Every step exists for a good reason. None is fast enough to beat a browser tab.

The result is a widening gap between what leadership believes is happening and what is happening. McKinsey's Superagency in the Workplace research (118 C-suite leaders and 3,002 employees, US, October-November 2024) found executives estimated only 4% of employees used generative AI for 30% or more of their daily work - while employees self-reported the true figure at 13%, more than three times higher. If that gap exists at the level of general daily use, it is reasonable to expect a comparable or larger gap around the ungoverned, undisclosed use that carries the most risk - precisely the use employees have the strongest incentive not to report.

What most institutions get wrong

The default response is a policy update: a revised acceptable-use clause, a reminder email, occasionally a network-level block on consumer AI domains. Three things are usually wrong with it.

It treats disclosure as the failure mode, when speed is the actual cause. A ban or a policy memo addresses the symptom (staff using unauthorised tools) without touching the cause (the authorised path being too slow to be worth using). Research on managers evaluating AI-assisted work from HEC Paris is instructive here, even outside financial services: when employees disclosed their AI use, managers rated the same work less favourably than when the AI involvement was undisclosed - a genuine disincentive to be honest that policy alone cannot undo. Tightening the policy without shortening the deployment cycle pushes disclosure down further; it does not reduce usage.

It assumes governance and enablement are in tension, when the ACA Group data suggests most firms have neither. With only 12% of AI-using financial firms holding a formal risk management framework and 92% lacking any third-party AI policy, the realistic starting point is not "loosen or tighten controls" - it is "we do not yet have controls that function at the speed staff already operate at."

It optimises for the wrong metric. Most AI governance programmes report on policy coverage, training completion and incident counts. None measure the variable that actually determines whether shadow AI persists: how long it takes, from the moment a genuine use case is identified, to put a governed and sanctioned equivalent in front of the people who need it.

The Desire Path Protocol

Urban planners have a term for the worn dirt trail that appears across a lawn when the paved path doesn't go where people actually need to walk: a desire path. It is not vandalism, and it is not a discipline problem - it is data about a design that does not match real behaviour. The right response was never a fence; it is to study where the paths already run and, where the need is legitimate, pave over them.

sourceCode applies the same logic to shadow AI through a four-stage diagnostic and response model: the Desire Path Protocol.

A four-stage diagnostic and response model for shadow AI in regulated back offices

1. Trace. Find the paths that already exist before writing anything new: pull DLP/CASB logs for traffic to consumer AI domains, run a short, genuinely no-penalty disclosure amnesty ("tell us what you use and why, no consequences"), and review team-level workflow retrospectives. Given KPMG's finding that over half of employees actively avoid disclosing AI use, a policy-attestation survey alone will systematically undercount - the trace needs behavioural signals, not just self-reports.

2. Grade. Classify each discovered use case on two axes: data sensitivity (does the input touch customer PII, financial data, material non-public information or regulated advice) and decision consequence (is the output an informational draft, or does it feed directly into a credit, claims, pricing or disclosure decision). This produces a small number of risk tiers rather than a single allow/block switch - most desire paths turn out to be low-sensitivity, low-consequence work like correspondence drafting or meeting summarisation, exactly the work that should be paved first.

3. Pave. For every high-frequency, tier-appropriate use case, commit to and publicly track a deployment SLA - elapsed time from "use case identified" to "governed, sanctioned equivalent live," built where possible on infrastructure already cleared (an existing enterprise LLM gateway, an already-approved model) rather than a bespoke build per request. This is the number that matters more than policy coverage: if the sanctioned path consistently beats the browser tab, staff have no reason to route around it.

4. Patrol. Desire paths regrow. Re-run the trace on a fixed cadence - quarterly is a reasonable default - and treat any reopened workaround around a "paved" capability as a design signal, not a compliance failure: the sanctioned tool likely has a latency, access or usability gap to fix, not a policy to restate.

The Protocol is deliberately not a ban-versus-allow decision. It is a way to make the deployment-speed variable visible and manageable, which is the variable every piece of evidence above points to as the actual driver of shadow AI.

Business and technology implications

Beyond the headline risk of sensitive data leaving the perimeter through a consumer tool's training pipeline or vendor logs, four implications are specific to regulated financial services:

- Audit trail gaps. Work created via a personal AI account typically leaves no retrievable record in the firm's systems of record - a material problem when a regulator, auditor or court later asks how a decision or customer communication was produced. Smarsh's UK data found 38% of financial services employees do not know whether their firm captures AI outputs at all.

- Model risk without a model risk process. A shadow AI tool used to draft a credit memo or summarise a claims file is, functionally, an unvalidated model in the decision chain - invisible to the model risk function by definition.

- Unassessed third-party exposure. Every consumer AI tool an employee uses is, in substance, an unassessed service provider handling firm and customer data. In Australia, APRA's CPS 230 (effective 1 July 2025) puts board-level accountability on operational risk and requires material service providers to meet equivalent risk standards - a requirement shadow AI routes around entirely, through invisibility rather than malice.

- Unsupervised quality judgement. Where no sanctioned tool exists, staff aren't only exposing data - they're making unreviewed calls about which consumer model to trust for financial-services-grade accuracy.

None of this argues for slower deployment. It argues for deployment that is fast and visible - the two are not in tension once speed is treated as a governance requirement rather than an engineering nice-to-have.

A fair counterargument: bans have worked, at least for a while

It would be dishonest to claim outright bans never work. In early 2023, JPMorgan Chase, Citigroup, Bank of America, Deutsche Bank, Goldman Sachs and Wells Fargo all restricted or banned employee use of public ChatGPT, and none suffered a visible operational collapse as a result. For a period, prohibition was a workable stopgap.

The more instructive detail is what happened next: most of these banks did not leave the ban in place indefinitely. They used the restriction window to stand up gated, enterprise-controlled versions of the same capability, then migrated staff onto those. Datos Insights analyst Gilles Ubaghs, speaking to American Banker, put it plainly: "there are always going to be workarounds" - the more durable response has been broader, governed rollout rather than permanent restriction. Zions Bank's approach, described by its CTO Jennifer Smith, follows the same pattern: enablement through approved portals and governance, on the view that "the benefits of generative AI outweigh the risks when managed effectively."

The honest reading of banking's own experience, then, is not "bans don't work." It is "bans buy time, and the institutions that used that time well spent it building the sanctioned alternative, not writing a longer policy." That is consistent with the Desire Path Protocol above - a ban can be the highest-risk-tier response while "Pave" is underway, provided it is a bridge and not the destination.

What leaders should do next

Four steps translate the above into something you can start this quarter, ahead of any formal AI strategy refresh:

1. Instrument before you legislate. Get read access to DLP/CASB traffic to known consumer AI domains before drafting a single new policy clause - you need the trace before you can grade anything accurately.

2. Run one disclosure amnesty. A genuinely no-penalty, time-boxed window for staff to declare current AI tool use will surface more real use cases in two weeks than a year of attestation-based audits, given how strongly the KPMG data shows concealment is already the norm.

3. Set and publish a deployment SLA, not just a policy. Pick the highest-frequency, lowest-risk-tier use case your trace surfaces, and commit publicly to an elapsed-time target for a sanctioned equivalent to go live. Hitting that target once, visibly, does more to change behaviour than a new clause in the handbook.

4. Route model risk and third-party risk functions into the same conversation as IT and business unit leads. Shadow AI sits at the intersection of all three; treating it as a pure security or pure compliance problem guarantees at least one blind spot.

The sourceCode’s perspective

We build and deploy technology for banks, insurers and fintechs across APAC and the Gulf, and the pattern above is consistent with what we see in delivery conversations: institutions most exposed to shadow AI are rarely the ones with weak governance intent. They are the ones where governance and delivery speed are treated as separate workstreams, owned by separate teams, on separate timelines. By the time a risk committee finishes evaluating a sanctioned tool, the underlying use case has often already migrated to a personal account months earlier - and stayed there, because nothing has beaten it since.

Our view is that the deployment-speed variable in the "Pave" stage is the one most institutions have never measured, let alone targeted, and the one a capable delivery partner can most directly move - not by lowering the governance bar, but by making the path from an identified, risk-graded use case to a live, audited, sanctioned equivalent short enough that it wins on its own merits.


That is a delivery capability question as much as a policy question, and it is worth raising before the next internal audit finds the gap for you. Talk to us!

Conclusion

The evidence - MIT NANDA, KPMG and The University of Melbourne, the US Federal Reserve, and financial-services-specific research from Smarsh and ACA Group - points in one direction: shadow AI in financial services back offices is not a compliance footnote, it is the current default state of AI adoption, running ahead of visibility, training and governance almost everywhere it has been measured. Banking's own history with the 2023 ChatGPT bans shows prohibition can buy time - but only the institutions that used that time to build a faster sanctioned path actually closed the gap.


The question worth putting to your leadership team this quarter is not whether staff are using unsanctioned AI tools for real work; evidence this consistent leaves little doubt that they are. The question is how long it currently takes your organisation to turn a discovered use case into a governed one - and whether that number is fast enough to make the workaround pointless.

Frequently Asked Questions

What is "shadow AI" in a financial services context? Shadow AI refers to employees using consumer-grade AI tools - such as personal ChatGPT, Claude or Gemini accounts - to complete real work tasks (drafting correspondence, summarising files, analysing documents) without the organisation's knowledge, approval, logging or governance. It differs from a failed official AI pilot in that it is usually invisible to IT and risk functions rather than a known, tracked initiative.

How common is shadow AI use in banks and insurers specifically? UK-specific research found 37% of financial services and insurance employees often use public AI tools for work (Smarsh, 2025). Global, cross-industry research from KPMG and The University of Melbourne (2025) found nearly half of employees admit to AI use that contravenes their organisation's policy, with over half concealing their use altogether - there is no reason to expect financial services is meaningfully below these broader averages given the sector-specific data available.

Isn't the simplest fix just to block consumer AI tools at the network level? Blocking can reduce visible usage on managed devices, and several major banks did this in 2023 without immediate operational disruption. But it does not address personal devices, personal accounts accessed outside the corporate network, or the underlying reason staff sought out the tool. The banks that blocked consumer AI in 2023 largely followed up with a sanctioned enterprise alternative rather than relying on the block indefinitely.

What is the single biggest risk shadow AI creates for a risk or compliance function? Loss of audit trail and model risk visibility. Work influenced by an unsanctioned AI tool typically leaves no retrievable record in the firm's systems of record, and the tool itself sits outside standard model risk and third-party risk assessment - making it, in effect, an unvalidated and unassessed input into regulated decisions.

Does APRA or other regulators specifically require oversight of employee AI tool use? Not yet in explicit terms. APRA's CPS 230 (effective 1 July 2025) requires board-level accountability for operational risk and equivalent risk standards from material service providers, but it does not currently name informal staff AI tool use directly. The exposure is best understood as an extension of existing operational and third-party risk obligations rather than a distinct, separately regulated category - for now.

How fast should a "sanctioned equivalent" realistically be deployed once a shadow AI use case is identified? There is no universal industry benchmark yet, which is itself part of the problem this article describes. The more useful question for most institutions is not "what is the right number" but "do we currently track this number at all" - most do not. sourceCode's Desire Path Protocol treats setting and publishing that target, however conservative it starts, as the more urgent first step.

Reference List

Challapally, A., Pease, C. and Raskar, R. (2025) The GenAI Divide: State of AI in Business 2025. MIT NANDA / Project NANDA. Available at: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf (Accessed: 23 September 2026).

KPMG International and The University of Melbourne (2025) Trust, Attitudes and Use of Artificial Intelligence: A Global Study 2025. Available at: https://assets.kpmg.com/content/dam/kpmgsites/xx/pdf/2025/05/trust-attitudes-and-use-of-ai-global-report.pdf (Accessed: 23 September 2026).

McKinsey & Company (2025) 'Leaders underestimate employees' AI use', Superagency in the Workplace. Available at: https://www.mckinsey.com/featured-insights/week-in-charts/leaders-underestimate-employees-ai-use (Accessed: 23 September 2026).

Bick, A., Blandin, A. and Deming, D. (2025) 'The Impact of Generative AI on Work Productivity', Federal Reserve Bank of St. Louis, On the Economy. Available at: https://www.stlouisfed.org/on-the-economy/2025/feb/impact-generative-ai-work-productivity (Accessed: 23 September 2026).

Smarsh (2025) New Smarsh UK Survey Shows AI Adoption Surges in Financial Services-But Employees Demand Stronger Guardrails. Fielded by OnePoll, 2,000 UK financial services and insurance employees, 8-17 April 2025. Available at: https://www.businesswire.com/news/home/20250625958908/en/New-Smarsh-UK-Survey-Shows-AI-Adoption-Surges-in-Financial-Services%E2%80%94But-Employees-Demand-Stronger-Guardrails (Accessed: 23 September 2026). [Note: commercially commissioned by Smarsh, a compliance/communications data vendor to regulated industries; cited for disclosed independent fieldwork methodology and sample size.]

ACA Group and National Society of Compliance Professionals (2024) 2024 AI Benchmarking Survey. Available at: https://www.acaglobal.com/news-and-announcements/financial-services-firms-lag-ai-governance-and-compliance-readiness-survey-reveals/ (Accessed: 23 September 2026). [Note: ACA Group is a compliance/GRC advisory firm with a commercial interest in governance findings; conducted jointly with the National Society of Compliance Professionals, a member-based professional association, and methodology and sample size are disclosed.]

American Banker (2025) 'Rise of shadow AI presents dilemmas for tech leaders'. Available at: https://www.americanbanker.com/news/rise-of-shadow-ai-presents-dilemmas-for-tech-leaders (Accessed: 23 September 2026).

American Banker (2025) 'Banks navigate workers' use of ChatGPT, set AI policies'. Available at: https://www.americanbanker.com/news/banks-navigate-workers-use-of-chatgpt-set-ai-policies (Accessed: 23 September 2026).

Clifford Chance (2025) 'Navigating Operational Risks: CPS 230's Influence on AI and Cybersecurity Strategies'. Available at: https://www.cliffordchance.com/insights/resources/blogs/regulatory-investigations-financial-crime-insights/2025/04/cps-230-influence-on-ai-and-cybersecurity-strategies.html (Accessed: 23 September 2026).

Related articles

10/09/2026

The Underwriting Data Problem Agentic AI Can't Fix

18/09/2026

The Real Difference Between a Claims Automation Pilot and a Claims Automation System

11/09/2026

Anatomy of a Real AI Vendor Exit Plan: The Seven Components Procurement and Risk Teams Actually Need

22/09/2026

What Underwriting Loses When It Optimises Only for Speed

16/09/2026

Embedded Insurance Is Growing Faster Than the Core Systems Behind It

15/09/2026

The Insourcing Decision Most CTOs Get Backwards: Capacity vs. Capability

17/09/2026

What Three Failed AI Vendor Selections Have in Common (A Procurement Post-Mortem)

21/09/2026

Why CPS 230 Changes How Australian Insurers Should Be Writing Technology Contracts

23/09/2026

The Hidden Cost of Shadow AI in Financial Services Back Offices

14/09/2026

Financial Crime Didn't Get Smarter. Transaction Monitoring Got Slower Relative to It

Navigating the Future of Software

linkedin
About usResources
SolutionssBrainChatbotVoicebotVoice RecognitionFace Recognition
Blog and InsightsAI & Blockchain Trends Industry Case Studies Thought Leadership Articles Success Stories & Client Spotlights 
Legal Privacy Policy Terms of Service 
linkedin

Australia - Malaysia - Vietnam

Copyright © 2026 source[code].

Australia - Malaysia - Vietnam