What Three Failed AI Vendor Selections Have in Common (A Procurement Post-Mortem)
Key Takeaways
-
MIT NANDA's 2025 research found 95% of enterprise generative AI pilots show no measurable profit-and-loss impact - but tools bought from specialised vendors succeeded roughly twice as often (66%) as tools built in-house (33%). The model is rarely the failure point.
-
Documented AI operating failures at a major Australian bank, a global fintech lender, and a US insurer - each independently reported and regulator-referenced - share a common root cause: the institution or vendor could demonstrate what the AI produced, but not the monitoring, escalation, or audit discipline needed to run it safely at scale.
-
Four financial regulators across three regions (APRA in Australia, MAS in Singapore, CBUAE in the UAE, and the NAIC in the US) have independently converged on the same finding: third-party AI accountability is the widest gap between current institutional practice and supervisory expectation.
-
Most procurement scorecards still weight demo performance and pricing far more heavily than operating discipline - creating a structural blind spot that a strong sales demo can walk straight through.
-
Not every AI vendor failure is a governance failure; genuine capability mismatch, unclear business value, and change-management resistance are real and separate causes. A good evaluation framework has to be able to tell the difference.

Introduction
Procurement teams evaluating AI vendors are working from an evidence base almost nobody has assembled properly. Every RFP for an AI capability draws on internal experience, a handful of analyst reports, and a vague sense that "everyone else is struggling with this too." What's missing is a structured way to learn from AI vendor selections that stalled or drew regulatory attention elsewhere in the industry, before committing budget and reputational capital to the next one.
That evidence base does exist, scattered across academic research, regulatory correspondence, and financial journalism rather than sitting in one place. MIT's NANDA initiative mapped the broader enterprise AI failure pattern in 2025. Four financial regulators, in Australia, Singapore, the UAE and the US, have each published guidance in the past eighteen months converging on the same specific gap. And a small number of AI-related operating failures in financial services have been documented in enough detail to see how the pattern plays out.
This is a post-mortem built from that evidence, not a stylised "three failures" narrative. It's worth being precise about what the public record contains: procurement failures - a vendor rejected at RFP stage, a contract that stalled in negotiation - are rarely disclosed, because nobody issues a press release about a deal that didn't happen. What is documented, repeatedly, is what happens next: AI systems that passed evaluation, entered production, and failed in ways that trace back to the same gap procurement should have tested for at selection. That is the more honest, and more useful, story to tell.
What The Evidence Actually Shows: Pilots Don't Usually Fail On The Model
The headline finding from MIT NANDA's "State of AI in Business 2025" report is stark: across the organisations studied, roughly $30-40 billion in enterprise generative AI investment has produced measurable profit-and-loss impact in only about 5% of cases. The researchers describe this as the "GenAI Divide" - a small number of organisations extracting real value while the large majority see pilots stall (MIT NANDA, 2025).

The more useful finding for procurement sits one layer down. AI tools acquired through purchase or strategic partnership with a specialised vendor reached full deployment roughly 66% of the time, versus about 33% for tools built internally - a two-to-one difference (MIT NANDA, 2025; Estrada, 2025). That is not an argument that vendors are the problem; if anything it's the opposite - buying beats building by a wide margin, on average. But a 66% success rate still means roughly one in three vendor-sourced deployments doesn't reach durable production use, even when the vendor is legitimate and the model works. NANDA points to "the learning gap" as the primary driver - tools that don't retain context, adapt to institutional workflows, or improve from feedback, regardless of how capable the underlying model is in isolation.
Gartner's research points the same way from a different angle. In June 2025, Gartner predicted more than 40% of agentic AI projects would be cancelled before the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls as the three leading causes (Gartner, 2025). Gartner also estimated that of the thousands of vendors marketing "agentic AI," only around 130 represent genuinely agentic functionality - the rest is largely existing chatbot, RPA, or assistant technology rebranded to match the moment, a practice the firm calls "agent washing" (Gartner, 2025).
Read together, these findings say something useful: whether a vendor's product can do the task is usually answerable in a well-run demo, and most shortlisted vendors can answer it credibly. What determines whether the deployment survives contact with production is different: does it adapt to how this institution works, is its cost and value trajectory understood before scale, and does the institution have the controls to run it safely once live.
Procurement's Actual Decision Problem
This creates a structural problem for the people running AI vendor evaluations. A typical enterprise AI RFP scores vendors against criteria weighted heavily toward what can be observed in a proof-of-concept: accuracy on a benchmark dataset, latency, integration effort, price. These are legitimate criteria. But they're also precisely what a well-resourced vendor can optimise for in a sales cycle, because a proof-of-concept is a controlled environment - curated dataset, no adversarial inputs, no real operational load.
What a proof-of-concept structurally cannot show is what happens eighteen months into production: how the model drifts as real customer data diverges from training assumptions, how an escalation actually reaches a human when the system is uncertain, whether the audit trail a regulator asks for six months after a disputed decision actually exists, and what happens if the institution needs to exit the contract under pressure. Those properties determine whether a deployment that passed its pilot survives its second year - and they are close to invisible in a standard RFP scoring matrix, because nobody has built a structured way to test for them before the contract is signed.
This is not hypothetical. It's the specific gap financial regulators, working independently of each other and of MIT NANDA, have identified as the widest space between current practice and supervisory expectation.
Four Regulators, One Converging Finding
Between May 2025 and February 2026, four financial regulators across three regions published guidance on AI risk management, none responding to each other or to MIT NANDA's research. They arrived at strikingly similar conclusions independently.
Singapore's MAS published good practices for AI model risk management on 1 May 2025: ongoing post-deployment monitoring against defined thresholds for robustness, stability, data quality and fairness; documented escalation procedures, with material production changes requiring control-function sign-off; and expanded audit-trail documentation covering data lineage, feature selection and fairness testing. For third-party AI specifically, MAS calls for independent compensatory testing of vendor claims, contractual audit rights and performance guarantees, and business-continuity backups (Monetary Authority of Singapore, 2025).
Australia's APRA went further in naming where the gap sits. In supervisory correspondence, it identified third-party and supply-chain risk as "the widest gap" between current AI governance practice and regulatory expectation among the banks and insurers it reviewed - AI embedded inside vendor platforms with opaque fourth-party dependencies, contracts frequently lacking audit rights or incident-notification timelines, and institutions with technically compliant contracts that simply weren't monitoring vendor performance against them. APRA also flagged concentration risk - single-provider dependency with no tested exit strategy - and set 1 July 2026 as the CPS 230 compliance deadline for existing AI supplier arrangements (APRA, cited in MinterEllison, 2025/2026).
The UAE Central Bank's responsible AI guidance for licensed financial institutions, published 11 February 2026, states the principle directly: "the responsibility for AI outcomes remains with LFIs, even where functions are outsourced." It requires documented due diligence on third-party and cloud AI vendors, periodic bias stress-testing, and a formal human-oversight model calibrated to risk level (Central Bank of the UAE, cited in Pinsent Masons, 2026).
In the US, the NAIC's Model Bulletin on AI use by insurers - adopted 2023, now being implemented state by state - builds "Third-Party Vendor Management" into its governance framework as one of seven required components, on the basis that insurers remain accountable for AI systems they didn't build, including demonstrating due diligence and contractual safeguards (NAIC Model Bulletin, cited in Buchanan Ingersoll & Rooney, 2025).
Four supervisors, four jurisdictions, one shared conclusion: the AI risk institutions manage worst is the risk embedded inside a vendor relationship, not the risk inside a model they built themselves.
What The Public Record Shows When This Gap Becomes Visible
Three documented cases in financial services show what happens when this gap is tested in production rather than in an RFP. None is a story about a vendor lying in a demo - in each, the underlying AI system did broadly what it was built to do. The failure was in the discipline around it.

Commonwealth Bank of Australia, the country's largest bank, announced in July 2025 that it would cut 45 customer service roles, attributing the redundancies to a new AI voicebot it said had reduced weekly call volumes by roughly 2,000. Within weeks the Finance Sector Union disputed the figures, reporting call volumes were in fact rising and team leaders were being pulled onto phones to cover the shortfall. On 21 August 2025 CBA reversed the decision and apologised, acknowledging it "did not adequately consider all relevant business considerations" and "should have been more thorough in [its] assessment of the roles required" (ABC News, 2025; Bloomberg, 2025). The bank did not dispute the technology functioned; the failure was that a workforce decision was executed on an AI performance claim that hadn't been independently validated against the operational data the union subsequently surfaced - precisely the production-monitoring step regulators now expect locked in before, not after, a decision of that consequence.
Klarna, the Swedish buy-now-pay-later lender, announced in February 2024 that an AI assistant built with OpenAI was handling the equivalent workload of 700 agents - 2.3 million conversations a month across 23 markets, cutting average handle time from 11 minutes to under two, with a projected $40 million profit improvement. By May 2025, CEO Sebastian Siemiatkowski acknowledged the quality had suffered: "cost unfortunately seems to have been a too predominant evaluation factor... what you end up having is lower quality." The mechanism is instructive: Klarna's tracked metrics - average handle time, aggregate satisfaction - looked healthy because they blended simple requests with complex ones. Degradation on higher-stakes cases (billing disputes, fraud reports, account closures) was masked by strong performance on routine queries; the monitoring design didn't separate the two (Bigeye, 2025; Entrepreneur, 2025). Klarna is now rehiring human agents. This wasn't a model failure - it was a measurement failure: a dashboard that couldn't see the thing that mattered.
GEICO, the US auto insurer, settled with the Pennsylvania Attorney General in May 2026 over an AI-enabled underwriting review tool that flagged a policyholder, then cancelled the policy "without adequate notice" when the customer's response was deemed insufficient - leaving the customer driving, uninsured, unaware. The settlement required GEICO to build a formal AI governance program with executive oversight, bias detection, and - explicitly - third-party vendor accountability, aligned to the NAIC's Model Bulletin (Clark Hill, 2026). The AI's underlying assessment may or may not have been defensible; the documented failure was procedural - a customer-facing automated decision with no adequate notice or human check built around it.
Three institutions, three technologies, three regulators - the same shape of failure: a system that could do what it was designed to do, operating without the monitoring, escalation or audit infrastructure to catch it when the real world diverged from the demo.
The Operating Discipline Scorecard: weighting what procurement can't see in a demo
The practical implication is that AI vendor evaluation criteria need to weight delivery and governance discipline at least as heavily as model capability - not as a compliance afterthought bolted onto the RFP, but as half the scoring weight from the outset. We call this the Operating Discipline Scorecard (ODS): a five-criterion evaluation structure that splits scoring evenly between what a vendor can demonstrate and what a vendor can prove it will sustain.

Capability (50% of score)
1. Model and data fit (25%) - accuracy, robustness, and relevance, tested against the institution's actual data, not a vendor benchmark.
2. Integration and adaptability (25%) - evidence the system absorbs institutional workflow and feedback over time, addressing NANDA's "learning gap" rather than assuming a strong demo generalises to production.
Operating discipline (50% of score)
3. Production monitoring and drift detection (15%) - can the vendor show, not describe, a live monitoring stack tracking accuracy, fairness and stability against thresholds, segmented by case complexity rather than blended into one aggregate metric - the failure mode Klarna's case illustrates.
4. Escalation and human-in-the-loop design (15%) - a documented, testable path from AI uncertainty to a human decision-maker, with defined mandatory-escalation thresholds.
5. Audit trail, exit, and concentration risk (20%) - decision-level documentation sufficient for regulatory reconstruction, contractual audit rights, and a genuinely tested exit or substitution plan - not a termination clause that has never been exercised.
Scoring capability alone can reward the best demo. Scoring discipline alone can reward the most bureaucratic vendor over a genuinely capable one - the counterargument below matters here. The ODS forces both questions onto the same page, at the same weight, before the contract is signed rather than after the first production incident.
Business and Technology Implications
For CTOs and CDOs, this weighting changes technical due diligence before the RFP is issued: requiring vendors to demonstrate monitoring and escalation infrastructure live, in a sandboxed version of the institution's own environment, rather than accepting a slide on "enterprise-grade governance." For risk and compliance, the evaluation itself should produce the documentation a supervisor would later ask for - model provenance, data lineage, escalation logs from the pilot - rather than generating that evidence retroactively. For CFOs and procurement, it means treating monitoring, audit and exit infrastructure as priced, contracted deliverables with service levels, not implied obligations buried in a master services agreement.
The commercial implication for vendors, source[code] included, is direct: capability alone is no longer sufficient to win regulated-industry AI work. Institutions applying this scrutiny aren't being difficult for its own sake - they're responding to supervisory guidance that already exists in writing, in four jurisdictions.
A Fair Counterargument: Not Every Failure Is A Governance Failure
It would be a mistake - and dishonest - to present every AI vendor disappointment as a governance failure in disguise. Gartner's top three reasons for agentic AI project cancellation are escalating costs, unclear business value, and inadequate risk controls - only the third is a discipline issue in the sense this article means. The first two are economic: the project wasn't worth what it cost, independent of how well-governed it was. NANDA similarly finds "unwillingness to adopt new tools," poor user experience, and lack of executive sponsorship among the most common barriers - organisational failures, not vendor discipline failures.
There's also a genuine capability-mismatch mode governance criteria won't catch: a vendor whose model is well-monitored and auditable but simply isn't accurate or fit-for-purpose for the use case. An evaluation that over-weights process maturity risks favouring large, bureaucratic incumbents practised at compliance paperwork over smaller, more capable specialists who haven't yet built a governance function to match - a bad outcome for buyers, and an ironic one for a framework meant to improve decision quality. This is why the ODS holds capability and discipline at equal weight rather than subordinating one to the other.
What Leaders Should Do Next
The immediate step isn't commissioning a new governance programme - it's taking the current AI vendor shortlist and re-scoring it against operating discipline criteria the existing RFP probably didn't ask about. For most institutions, that means returning to the vendors furthest along in an active evaluation and requesting three artefacts before proceeding: a live demonstration of production monitoring (not a description), a documented escalation pathway with defined thresholds, and evidence the exit or substitution clause in the draft contract has actually been tested. Where a vendor can't produce those credibly, that's the answer the evaluation was designed to surface.
The source[code] Perspective
Our view, built from delivery work across banking, insurance, and fintech clients in Australia, Southeast Asia and the Gulf, is that the vendors best positioned to win regulated AI work over the next two years will not be the ones with the most impressive model demo. They will be the ones who can walk a procurement panel through their monitoring dashboard, escalation runbook, and audit log - live, on a real deployment, under questioning - with the same fluency they bring to a capability pitch. That is a different skill from building a good model, and the public record and four independent regulators say it's the one that matters most once the contract is signed.
Evaluation criteria should catch up to that reality before the next production incident does it for them. Talk to us!
Conclusion
The evidence assembled here - MIT NANDA's research, Gartner's, and the converging guidance of four financial regulators, cross-checked against three independently documented operating failures in banking, fintech lending and insurance - points to one consistent conclusion.
AI vendor selection in financial services fails less often because the model can't do the job, and more often because nobody tested, before signing, whether the institution and the vendor together had the operating discipline to run it safely once live.
Procurement teams that build that test into the evaluation now are buying insurance against the specific failure mode the industry has already, repeatedly, documented in public.
Frequently Asked Questions
What is the "GenAI Divide" and why does it matter for vendor selection? It's MIT NANDA's term for the gap between the small share of enterprise generative AI deployments (around 5%) that show measurable profit-and-loss impact and the large majority that don't. It matters for vendor selection because the research found buying from a specialised vendor succeeds roughly twice as often as building internally - but a third of vendor-sourced deployments still fail, meaning vendor selection quality, not the buy-versus-build decision alone, is doing real work.
Were the Commonwealth Bank, Klarna, and GEICO cases actual "vendor selection failures"? Not in the literal RFP sense - none is documented as a vendor being rejected or a contract collapsing during procurement. Each is a documented operating failure that occurred after an AI system was already selected and deployed, and each traces back to a gap - inadequate validation of performance claims, monitoring that masked degraded performance, and insufficient escalation around a customer-facing automated decision - that a discipline-weighted evaluation at selection stage is designed to catch before it reaches production.
How much weight should governance and monitoring capability really get in an AI vendor RFP? The framework in this article recommends an even split - 50% model and integration capability, 50% operating discipline (monitoring, escalation, audit trail, exit planning) - on the basis that regulatory guidance in Australia, Singapore, the UAE, and the US all identify third-party AI discipline as the most under-managed risk category, not a secondary consideration.
Does weighting governance this heavily risk favouring large incumbent vendors over more capable smaller ones? It's a real risk if discipline criteria are scored on documentation alone. The safeguard is requiring live demonstration - a working monitoring dashboard, a tested escalation path, an exit clause that has actually been exercised - rather than accepting policy documents, which specifically favours vendors who have operationalised discipline over those who have merely written about it.
Is every AI deployment failure actually a governance failure? No. Gartner and MIT NANDA both identify cost overruns, unclear business value, and organisational change resistance as common, separate failure causes unrelated to vendor governance. A sound evaluation framework needs to distinguish a genuine capability or value mismatch from a discipline gap, rather than treating every failure as evidence of poor governance.
What should a procurement team do this quarter if they're mid-evaluation on an AI vendor? Request three specific, demonstrable artefacts from the shortlisted vendors before signing: a live view of production monitoring (not a description), a documented escalation pathway with defined trigger thresholds, and evidence the contract's exit or substitution clause has been tested rather than merely drafted.
Reference List
APRA (Australian Prudential Regulation Authority), cited in MinterEllison (2025/2026) APRA's AI letter: A wake-up call for managing your third-party suppliers. Available at: https://www.minterellison.com/articles/apra-ai-letter-third-party-suppliers (Accessed: 17 September 2026).
Bigeye (2025) Klarna's AI customer service deployment | AI Autopsy 002. Available at: https://www.bigeye.com/blog/klarnas-ai-customer-service-deployment (Accessed: 17 September 2026).
Bloomberg (2025) Australia's Biggest Bank Reverses Plan to Replace Jobs With AI. 21 August. Available at: https://www.bloomberg.com/news/articles/2025-08-21/commonwealth-bank-reverses-job-cuts-decision-over-ai-chatbots (Accessed: 17 September 2026).
Buchanan Ingersoll & Rooney PC (2025) When Algorithms Underwrite: Insurance Regulators Demanding Explainable AI Systems. Available at: https://www.bipc.com/when-algorithms-underwrite-insurance-regulators-demanding-explainable-ai-systems (Accessed: 17 September 2026).
Central Bank of the UAE, cited in Pinsent Masons (2026) UAE Central Bank publishes responsible AI guidance for financial sector. 11 February. Available at: https://www.pinsentmasons.com/out-law/news/uae-central-bank-responsible-ai-guidance-financial-sector (Accessed: 17 September 2026).
Clark Hill (2026) GEICO AI Settlement Signals Insurance Compliance Risks. Available at: https://www.clarkhill.com/news-events/news/geico-ai-settlement-insurance-underwriting-compliance/ (Accessed: 17 September 2026).
Entrepreneur (2025) Klarna Is Hiring Customer Service Agents After AI Couldn't Cut It on Calls, According to the Company's CEO. Available at: https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396 (Accessed: 17 September 2026).
Estrada, S. (2025) 'MIT report: 95% of generative AI pilots at companies are failing', Fortune, 18 August. Available at: https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo (Accessed: 17 September 2026).
Gartner (2025) Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. 25 June. Available at: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 (Accessed: 17 September 2026).
MIT NANDA (2025) The GenAI Divide: State of AI in Business 2025. Challapally, A., Pease, C. and Raskar, R. Available at: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf (Accessed: 17 September 2026).
Monetary Authority of Singapore, cited in Clifford Chance (2025) Good Practices for AI Model Risk Management for Singapore Financial Institutions. 1 May. Available at: https://www.cliffordchance.com/insights/resources/blogs/talking-tech/en/articles/2025/05/good-practices-for-ai-model-risk-management-for-singapore-financ.html (Accessed: 17 September 2026).
National Association of Insurance Commissioners (NAIC) (2023) Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, cited in Buchanan Ingersoll & Rooney PC (2025), as above.
ABC News (2025) Commonwealth Bank backtracks on AI job cuts, apologises for 'error' as call volumes rise. 21 August. Available at: https://www.abc.net.au/news/2025-08-21/cba-backtracks-on-ai-job-cuts-as-chatbot-lifts-call-volumes/105679492 (Accessed: 17 September 2026).