AI

AI adoption has outrun understanding

Artificial intelligence is now mainstream in UK financial services, but the firms that deploy it have outrun the firms that understand it. A maturity assessment shows a risk function where it stands, and what to fix first.

A board asks the head of model risk a deceptively simple question: how much of our decisioning is AI-driven? It sounds like it should resolve to a single percentage. It does, but not the one the board expects. The honest reply is a measure most firms have never taken: not how many models use machine learning, but how much of the AI already in production the team can explain. For most UK firms that second number is far smaller than the first. The distance between them is what a maturity assessment surfaces.

That distance has a name. AI maturity is the capability to capture value from artificial intelligence at scale while keeping the risks it introduces under control, and a maturity assessment is the structured diagnostic that scores a function against it, conventionally across five dimensions: strategy, data, governance, talent, and technology. For a risk function the exercise isn’t a technology audit. It’s a reading of whether the team can develop, validate, and monitor AI-based models at the speed the business is adopting them, and whether it can answer for their behaviour when a supervisor asks.

The question has become urgent because the gap it measures is widening. Adoption has crossed from experiment to infrastructure, depth of understanding hasn’t kept pace, and the distance between the leading firms and the rest is growing. A maturity assessment is how a function turns that vague unease into a position it can defend.

What the surveys show

1

Bank of England and FCA

Artificial intelligence in UK financial services 2024

View source ↗

The clearest evidence of the gap comes from the joint surveys the Bank of England (BoE) and the Financial Conduct Authority (FCA) have run on AI in UK financial services. In 2022, 58% of regulated firms reported using or developing machine learning applications. The 2024 survey drew 118 responses from banks, insurers, and financial market infrastructure providers. By then adoption had reached 75%, with a further 10% planning it, though on a broader definition: the 2022 survey measured machine learning, the 2024 survey all AI, so part of the rise is scope rather than uptake.1

A new category had also appeared: foundation models, the large general-purpose systems behind generative AI, already accounted for 17% of reported use cases.

Governance scaffolding has broadly kept up: 84% of firms had a designated accountable person for AI by 2024. Comprehension hasn’t. Only 34% reported a complete understanding of the AI they deploy, and the shortfall was concentrated in third-party models, where the firm can’t inspect the training data or the methodology behind a vendor’s product.

75% / 34%Share of UK firms that deploy AI, against the share that fully understand what they deploy (BoE and FCA, 2024)

The divergence between what firms deploy and what they understand is the headline finding for any risk function. It means that for roughly two-thirds of UK firms, AI is in production while the firm itself can’t independently characterise how it behaves. Maturity isn’t the adoption number. It’s the second number, and closing the distance between the two is the work.

Which frameworks score AI maturity?

A maturity assessment needs a yardstick, and three are in common use.

2

BCG

Build for the Future 2025

View source ↗

The most detailed is the Build for the Future study from Boston Consulting Group (BCG), which in 2025 assessed 1,250 organisations against 41 maturity dimensions and sorted them into four bands, from stagnating to future-built. What separates the top band isn’t access to technology, which everyone now has. It’s how the firm is wired: the future-built embed AI into core workflows rather than isolated pilots, and reinvest the gains. The study puts the payoff at 1.7 times the revenue growth of their peers, and 3.6 times the total shareholder return.2

The second yardstick, from MIT Sloan Management Review and BCG, classifies firms by how deeply they learn from what they deploy. Passives have no meaningful AI activity. Experimenters run pilots but never scale them. Investigators deploy in selected areas without learning from what they run, and Pioneers deploy broadly and turn that learning into improvement. For a risk function the step that matters is Investigator to Pioneer: the moment AI stops being a set of disconnected tools and becomes integral to how models are developed, validated, and monitored.

The third yardstick is the empirical one: the BoE and FCA survey levels, which tell a UK firm where the sector actually sits rather than where a framework says it could. That’s the yardstick the rest of this piece leans on, scored across the five dimensions the frameworks share, because for a risk function the question is never abstract capability, it’s position against the peers a supervisor also reads.

Four maturity bands: 60% sit in the bottom two

What are the five dimensions of AI maturity?

Whichever yardstick a function picks, the same five dimensions recur as the pillars behind the distance between what a firm deploys and what it understands, and each takes a specific meaning inside a risk function.

Strategy asks whether the function has a multi-year AI ambition tied to concrete risk outcomes, faster model development, wider validation coverage, lower monitoring latency, and whether senior leadership owns it. BCG names the lack of top-management commitment as one of the biggest reasons firms fall behind, and it’s the understanding gap that goes unresourced when nobody senior owns the ambition.

Data covers the availability, quality, and integration of the data needed to train and monitor AI models. Explainability starts here: a team can’t characterise a model’s behaviour if it can’t trace what the model was trained on and what it is fed in production.

In practice

The five dimensions are interdependent, not a menu to pick from. A function can have a board-backed strategy and strong talent, but if its AI runs on third-party models it can’t inspect, its governance score caps the others. Maturity is set by the weakest load-bearing dimension, not the average across all five.

3

Gini

SS1/23: model risk management for UK banks

View source ↗
4

PRA

SS1/23: Model risk management principles for banks

View source ↗

Governance anchors AI to the firm’s model risk framework, in the UK the Prudential Regulation Authority’s Supervisory Statement SS1/23, in force since 17 May 2024. Our companion piece on SS1/23 covers the full detail3. Principle 1 asks for “an established definition of a model that sets the scope for MRM, a model inventory and a risk-based tiering approach”; Principle 2 asks the board to “appoint an accountable individual to assume the responsibility to implement a sound MRM framework”, with independent validation under Principle 4.4 It’s the dimension that decides whether the firm can answer for the AI it runs.

Talent is the people who do the explaining. Validation teams increasingly need data and software engineering skills next to traditional statistics, because the model estate grows faster than headcount, and a reviewer who can’t read a vendor’s methodology can’t challenge it.

Technology covers the platforms that move a model from a notebook to monitored production. The test is simple: a scorecard running as a script on one analyst’s laptop is a maturity gap. The same scorecard versioned, deployed, and feeding its decisions to a dashboard the validation team can interrogate is maturity itself.

Five dimensions, only as strong as the weakest

Where the rulebook stops

Governance is the dimension where AI maturity most often breaks, and the reason is structural. SS1/23 is the UK’s main model-risk rulebook. Its definition turns on whether a method applies statistical, economic, financial or mathematical theory to process input data into output, which reaches a machine learning model comfortably. It says nothing at all about generative or agentic systems: the words appear nowhere in the statement.4 That is silence rather than a carve-out, and it leaves the question open. The exclusion matters because generative AI is the fastest-growing part of the estate. Agentic systems are the ones that work through a task in steps, choosing their own route and tools as they go.

17%Foundation models’ share of UK AI use cases in 2024, a category absent in 2022 (BoE and FCA, 2024)

The conventional validation a risk function relies on assumes a model with a fixed set of inputs and a deterministic output: feed it the same case and it returns the same answer, so a finite test set can characterise it. Agentic AI breaks that assumption. The same prompt can produce different reasoning steps and different tool calls depending on the system’s state, so a test set can’t cover the range of production inputs, and there is no single output function to certify.

The consequence for maturity is that governance can’t stop at pre-deployment sign-off. Industry frameworks increasingly argue it has to extend to continuous oversight: structured logging of inputs, reasoning steps and tool calls as a precondition for deployment rather than optional observability, and human checkpoints calibrated to the materiality of each decision rather than applied as blanket review. None of that is binding, which is precisely the gap.

A mature risk function, then, isn’t one that has validated its AI once. It’s one that has built the instrumentation to keep watching it, particularly for the generative and third-party models that the rulebook does not address and that sit inside the two-thirds understanding gap.

Validating agentic AI: one path you can test, or a space you cannot

What supervisors now assume

5

FSB

The Financial Stability Implications of Artificial Intelligence

View source ↗
6

OECD

Regulatory approaches to artificial intelligence in finance

View source ↗

The supervisory view reinforces the same point from the outside. The Financial Stability Board’s 2024 update to its 2017 report on AI in finance finds four vulnerabilities that “stand out for their potential to increase systemic risk”, among them concentration among a small set of third-party AI providers and the limited explainability of the models firms are buying.5 The Organisation for Economic Co-operation and Development’s 2024 survey of financial supervisors found most of them extending existing model-risk principles to AI rather than writing supervisory AI rules, which is the UK’s position.6 The EU has done both: its supervisors extend existing frameworks, while the EU AI Act classifies creditworthiness assessment as high-risk under Annex III, with those obligations phasing in through 2026 and 2027.

Neither body prescribes a maturity model. But both presume something that should concentrate the mind: that a supervised firm can assess its own AI capability against a structured taxonomy, and that an examiner can verify the assessment. Neither says how.

The gap worth closing first

For a risk function, the practical value of a maturity assessment is that it converts a vague sense of being behind into a specific, rankable list of gaps. The frameworks differ in their detail, but they agree on where to look first, and so does the UK evidence: the distance between what a firm deploys and what it understands.

The action is concrete and immediate. Map the AI already running in the function against the five dimensions, and for each model record one fact above the others: whether the team can independently characterise its behaviour, or whether it’s trusted on the vendor’s word. The models that fail that test, the third-party systems and the generative tools SS1/23 says nothing about, are where maturity is genuinely thin, however high the headline adoption figure looks. Closing that gap is unglamorous work: validation coverage, monitoring instrumentation, documentation, and the people to do it. But it’s the work that turns an impressive deployment statistic into a capability a board can stand behind and a supervisor can test.

The next time the board asks how much of the decisioning is AI-driven, the answer that counts is the share the team can explain.

Frequently asked questions

What is AI maturity, and what does an assessment measure?

AI maturity is the capability to capture value from artificial intelligence at scale while keeping the risks it introduces under control, and a maturity assessment is the structured diagnostic that scores a function against it, conventionally across five dimensions: strategy, data, governance, talent and technology. For a risk function it is not a technology audit. It reads whether the team can develop, validate and monitor AI-based models at the speed the business is adopting them, and whether it can answer for their behaviour when a supervisor asks.

What do the Bank of England and FCA surveys show about AI adoption?

The 2024 joint survey drew 118 responses from banks, insurers and financial market infrastructure providers, and found 75% of respondents using AI with a further 10% planning it. Foundation models, the large general-purpose systems behind generative AI, already accounted for 17% of reported use cases, a category absent from the previous survey. The 2022 survey reported 58% using or developing machine learning applications, so the two figures are not directly comparable: the earlier one measured machine learning and the later one measures AI on a broader basis.

What is the understanding gap?

Governance scaffolding has broadly kept pace with adoption while comprehension has not. By 2024, 84% of firms had a designated accountable person for AI, but only 34% reported a complete understanding of the AI they deploy. The shortfall concentrated in third-party models, where a firm cannot inspect the training data or the methodology behind a vendor's product. For around two-thirds of UK firms, then, AI is in production while the firm cannot independently characterise how it behaves.

What are the five AI maturity dimensions for a risk function?

Strategy asks whether the function has a multi-year AI ambition tied to concrete risk outcomes, and whether senior leadership owns it. Data covers the availability, quality and integration of the data needed to train and monitor models, and explainability starts here, because a team cannot characterise a model's behaviour without tracing what it was trained on and what it is fed in production. Governance anchors AI to the firm's model risk framework. Talent is the people who do the explaining, increasingly needing data and software engineering skills alongside statistics. Technology covers the platforms that move a model from a notebook to monitored production.

Why is maturity set by the weakest dimension rather than the average?

Because the five are interdependent rather than a menu to pick from. A function can have a board-backed strategy and strong talent, and still be capped by governance if its AI runs on third-party models it cannot inspect. Averaging across the five produces a flattering score that conceals the load-bearing weakness, which is the one a supervisor will find.

Which AI maturity frameworks are in common use?

Three. Boston Consulting Group's Build for the Future study assessed 1,250 organisations against 41 maturity dimensions in 2025 and sorted them into four bands: stagnating at 14%, emerging at 46%, scaling at 35%, and future-built at 5%, so 60% sit in the bottom two. What separates the top band is not access to technology, which everyone now has, but whether AI is embedded in core workflows rather than isolated pilots. The MIT Sloan Management Review and BCG classification sorts firms by how deeply they learn from what they deploy, from Passives through Experimenters and Investigators to Pioneers. The third yardstick is empirical: the Bank of England and FCA survey levels, which show where the UK sector actually sits.

Why does governance break first?

Because the rulebook's edge falls in an awkward place. SS1/23, in force since 17 May 2024, is the UK's main model risk statement, requiring a firm-wide model inventory, risk-based tiering, independent validation and a named senior manager accountable for the framework, and it reaches machine learning models through its materiality-based definition. How far it reaches generative and agentic AI is less settled, and the boundary is not drawn as explicitly on the face of the statement as commentary sometimes suggests. Since generative AI is the fastest-growing part of the estate, a firm that waits for the boundary to be clarified is accumulating models nobody has decided how to govern.

Why can't conventional validation handle agentic AI?

Conventional validation assumes a model with a fixed set of inputs and a deterministic output: feed it the same case and it returns the same answer, so a finite test set can characterise it. Agentic systems work through a task in steps, choosing their own route and tools as they go, so the same prompt can produce different reasoning steps and different tool calls depending on the system's state. A test set cannot cover the range of production inputs, and there is no single output function to certify.

What does governance for agentic AI have to add?

It cannot stop at pre-deployment sign-off. Emerging practice extends it to continuous oversight: structured logging of inputs, reasoning steps and tool calls treated as a precondition for deployment rather than optional observability, and human checkpoints calibrated to the materiality of each decision rather than applied as blanket review. A mature risk function is not one that validated its AI once, but one that built the instrumentation to keep watching it.

What do international bodies say about AI in finance?

The Financial Stability Board's 2024 update to its 2017 report finds four vulnerabilities that “stand out for their potential to increase systemic risk”, among them concentration among a small set of third-party AI providers, and the limited explainability of the models firms are buying. The Organisation for Economic Co-operation and Development, surveying regulatory approaches in 2024, found supervisors mostly extending existing model risk principles to AI rather than writing supervisory AI rules, which is the UK's position; the EU does both, since the EU AI Act treats creditworthiness assessment as high-risk with obligations phasing in through 2026 and 2027. Neither body prescribes a maturity model, but both presume a supervised firm can assess its own AI capability against a structured taxonomy and that an examiner can verify the assessment.

Where should a risk function start?

By mapping the AI already running in the function against the five dimensions, and recording one fact for each model above the others: whether the team can independently characterise its behaviour, or whether it is trusted on the vendor's word. The models that fail that test, the third-party systems and the generative tools whose governance status is unsettled, are where maturity is genuinely thin however high the headline adoption figure looks. Closing the gap is unglamorous: validation coverage, monitoring instrumentation, documentation, and the people to do it.

Sources

  1. 1 Bank of England and FCA. Artificial intelligence in UK financial services 2024 View source ↗
  2. 2 BCG. Build for the Future 2025 View source ↗
  3. 3 Gini. SS1/23: model risk management for UK banks View source ↗
  4. 4 PRA. SS1/23: Model risk management principles for banks View source ↗
  5. 5 FSB. The Financial Stability Implications of Artificial Intelligence View source ↗
  6. 6 OECD. Regulatory approaches to artificial intelligence in finance View source ↗
Receive updates directly in your inbox

Stay connected