Banking

The TrueVision transition: bridging from the old score to the new

How to swap a bureau score for a new one without redeveloping every model that depends on it.

Replacing a bureau score is a model risk exercise, not just an IT project. The task is to keep the models built on the old score producing reliable predictions, not simply to switch a data feed.

1

PRA SS1/23

Model risk management principles for banks

View source ↗
2

EBA/RTS/2026/05

material model changes, amending Delegated Regulation (EU) No 529/2014

View source ↗

Both the UK and European regimes treat replacing a bureau score as a model change. The Prudential Regulation Authority (PRA) made SS1/23 technology-agnostic and applied it across the full model lifecycle, so the obligation attaches when the feed changes rather than when a model is eventually rebuilt.1 And for an Internal Ratings Based (IRB) bank the classification test is specific. In March 2026 the European Banking Authority (EBA) issued final draft regulatory technical standards (RTS) on material model changes. These put the inclusion or removal of risk drivers in the notification tier, except where a change moves "the rank ordering or the distribution of obligors or exposures across grades or pools, in a fundamental manner, the measure and level of which will have been defined by the institution".2 A score substitution creates exactly that case.

That was the argument of our earlier article on the TrueVision transition; this piece goes inside the mechanics of the swap.

The short-term answer for most lenders is a bridge, not a rebuild. The transition eventually leads to a full redevelopment of every downstream model, but the immediate task is to keep those models producing the same decisions with the new score in place.

This article focuses on one specific pattern: the bureau score fed directly into the calibration layer that produces the model’s output. A calibration layer is the piece of the model that turns a score into a probability. Two other patterns exist. Firstly, the bureau score can sit inside a wider scorecard alongside other variables, in which case the same bridging idea applies but a scorecard-level refit is required behind it. Secondly, some lenders (typically Internal Ratings Based banks and larger portfolios) build a bespoke internal model against the raw bureau data rather than a bureau score.

Adjusting for a different distribution

A calibration layer does not read a bureau score in isolation. It reads the particular distribution the score had at the time the model was built or fitted. TrueVision has no reason to share the old score’s scale or shape, so the parameters that weight the score now operate on a different quantity, and the probabilities shift.

TRUEVISION · SCORE DISTRIBUTION: Same borrowers, re-scored

Downstream of the model, the same old score is written into decision rules that need to be revisited: the origination cut-off may have been set at a value that admitted a defined share of applicants, and pricing bands were drawn to place borrowers across risk tiers in known proportions. Every one of these numbers depends silently on the shape of the old score, so once TrueVision is substituted in, the cut-off no longer admits the share of applicants it was set to admit. Nothing fails on the day of the change. The failure is quiet, and it shows up later in approval rates and credit losses.

The lender has two options for bridging to the new score, and the data usually dictates which is right. If TrueVision scores can be paired with outcomes for a historical sample, the calibration can be refitted from scratch. If only the two scores overlap on the current book, without paired outcomes, quantile mapping is the alternative.

Fresh score-to-outcome calibration

The proper answer, when the data supports it.

The calibration layer is a formula that maps a score value into the log-odds space of an account outcome. The shape of that mapping is controlled by two parameters: an intercept α (alpha) and a slope β (beta). Refitting them against the lender’s own outcomes is the most durable answer to a score swap.

This first path (Path A) requires TrueVision scores overlapping with the sample used to fit the current calibration. Either the new score has been live long enough for that overlap to have accumulated, or the lender purchases a retro TrueVision dataset from TransUnion, a historical rescoring of borrowers whose outcomes are already known.

With that data in hand, the modelling team refits α and β via logistic regression, fitting the new score directly against the lender’s outcomes. The calibration layer becomes native to TrueVision, a like-for-like replacement of the old score system.

PATH A · FRESH CALIBRATION: Fit α and β to the lender's own outcomes

Score-to-outcome recalibration is a routine refresh that lenders already run periodically. Lenders can run it more often through the transition period, as the new score’s performance history accumulates.

3

Butaru, Chen, Clark, Das, Lo and Siddique (2016)

Risk and risk management in the credit card industry

View source ↗

An 800 on TrueVision at one bank does not mean the same outcome rate as an 800 at another bank, because the two portfolios differ; each bank’s α and β have to reflect its own book. That is not just intuition. One study applied machine learning to account-level credit card data from six major commercial banks, covering January 2009 to December 2013. It found that risk factors, their sensitivities and the predictability of delinquency all varied significantly between institutions holding structurally comparable revolving portfolios, and concluded that risk parameters should be calibrated to each institution’s own exposures rather than to an industry benchmark.3 A fresh calibration puts the mapping on the lender’s own outcomes rather than on any bridging assumption.

Quantile mapping

The answer when the data is not there for Path A.

Path B requires only overlapping old-and-new score data on the lender’s book: a window of borrowers scored under both. No outcome data is needed. Without that overlap window, the path is not available.

The mechanism is a rank-preserving transformation, known formally as quantile mapping.

PATH B · QUANTILE MAPPING: Up, across, down: same rank, new scale

Wrap each borrower’s new score onto the old score’s scale at the same rank position, and feed the wrapped value into the existing calibration. If a borrower’s TrueVision score sits at the 80th percentile of the new score’s distribution, the calibration is fed the old-score value from the 80th percentile of the old distribution. The calibration then reads a value on the scale it was fitted against, and continues to operate.

Two things carry across.

  • The ranking survives. The mapping never reorders borrowers; discriminatory power carries through untouched.
  • The scale survives. The mapped values live on the old score’s distribution, so the calibration receives an input with the range and shape it was fitted against.

Whether the mapping holds up in production comes down to how the reference distributions are chosen. The mapping depends on a specific reference period, and a stale sample carries stale patterns into every model that uses the new score. The tie-breaking and tail-capping rules that make this work in practice are a job for the modelling team, not a decision this article needs to walk through.

Monitoring the bridge

Once the bridge to the new score is built and recalibrated, it still has to be monitored, because the population crossing it changes over time. The mix of applicants shifts as the economy and the lender’s own marketing change, so a score that looked stable at launch can drift.

Population stability index (PSI)

The standard way to measure that drift is through the population stability index (PSI). It splits the score into bands and compares the share of the population in each band today against the share in a chosen reference sample. Identical distributions score zero; the bigger the gap, the higher the index.

POPULATION STABILITY INDEX: Five bands, then a shift
4

Djurovic

Representativeness testing and the classifier two-sample test

View source ↗

One limitation is worth knowing before relying on the index: it is univariate, so it cannot detect a shift in the joint distribution that leaves each feature’s own distribution intact, and aggregating several features’ indices into a single number has no statistical interpretation.4 With that said, it is typical to interpret the PSI against three loose thresholds. The bands below are convention rather than rule, and the middle cut-off is drawn at 0.20 or 0.25 depending on house practice:

  • Below 0.10, the population is effectively unchanged.
  • Between 0.10 and 0.20, the drift merits a warning and closer monitoring.
  • Above 0.20, the model comes under review.

The last decision is which reference sample from the population to compare against. Two choices are useful, in different ways.

  • The original build sample. This shows how far the population has moved since the model was first built, which mixes the effect of the switch with years of ordinary drift.
  • The last sample before the switch. This isolates the effect of the switch itself, measured against a recent starting point.

Testing the swap

Stability monitoring tells you whether the population is moving. It does not tell you whether the swap itself is making the right calls. Two adjacent techniques help: a shadow engine runs the new model in parallel with the old model on the same live traffic, recording the decision each would have made without acting on the new one; an A/B split does the same by directing part of the traffic to each engine and acting on both. Either produces two comparison populations worth looking at.

  1. Accounts the old engine accepted but the new one would have rejected. These have observed outcomes and give the cleanest read on the value of the switch.
  2. Accounts the old engine rejected but the new one would accept. These have to be inferred. The mechanics of that inference, including reject-inference, are a topic in their own right and belong outside this article.

The right path, given the data

The choice between the two paths comes down to what the lender has on the book at the point of the switch. Two situations cover the field:

  1. Overlap only. The new and old scores are both available for a window of borrowers, but no paired outcomes on the new score. Path B (quantile mapping) is the answer: substitute the input by mapping the new score onto the old scale, realign the cut-offs and pricing bands the switch will have moved, and stand up PSI monitoring alongside a shadow engine or A/B comparison.
  2. Overlap and outcomes. Either a TrueVision retro dataset from TransUnion is available, or the new score has been live long enough for outcomes to accumulate on the lender’s own book. Path A (fresh calibration) is the answer: refit α and β against the lender’s outcomes, then realign cut-offs and stand up monitoring as above.

For lenders whose data will remain overlap-only, quantile mapping is the answer, full stop. For those who begin overlap-only but expect to acquire outcome data (via a retro purchase or accumulated live history), the two techniques chain: Path B first, Path A once the data supports it.

The transition is a real piece of work, but every technique here is standard equipment, familiar to any validation team. Built from a well-documented quantile map or a fresh calibration, a bridge will carry the book across intact and keep the models working. Understanding the underlying distributions is what separates a bridge that holds from one that has to be repaired in production.

In our experience helping lenders through comparable score transitions, the firms that plan the work and document it as they go are the ones whose regulators and boards see no blip in performance, and whose portfolios do not produce one.

Frequently asked questions

Is replacing a bureau score an IT change or a model change?

A model change, and the framework treats it as one. The score is a predictor inside every model fitted against it, so swapping it changes what those models read. The EBA's March 2026 final draft RTS on material model changes puts the inclusion or removal of risk drivers in the notification tier, except where a change moves "the rank ordering or the distribution of obligors or exposures across grades or pools, in a fundamental manner, the measure and level of which will have been defined by the institution", which is precisely what substituting one bureau score for another does. Under the PRA's SS1/23 the whole lifecycle is in scope regardless of technique, so the documentation obligation starts at the point the feed changes, not at the point a model is rebuilt.

Should a lender bridge or rebuild?

Bridge first, rebuild eventually. The migration does end in a full redevelopment of every downstream model, but that is not the immediate task. The immediate task is keeping the existing models producing reliable decisions with a new score in place, and a bridge does that in weeks where a redevelopment programme takes quarters. Treating the bridge as the whole answer is the failure mode; treating a rebuild as the only answer is the delay.

What actually breaks on the day of the switch?

Nothing visible, which is what makes it dangerous. A calibration layer does not read a score in isolation, it reads the distribution that score had when the model was fitted. Substitute a score with a different scale and shape and the parameters now weight a different quantity, so the output probabilities shift. Downstream, the origination cut-off was set to admit a defined share of applicants and the pricing bands were drawn to place borrowers across tiers in known proportions. Every one of those numbers depends silently on the old score's shape. The model does not fail on the day. It shows up later in approval rates and credit losses.

What decides which bridging method to use?

The data on the book at the point of the switch, not a methodological preference. If new scores can be paired with known outcomes, whether because the score has been live long enough or because the lender buys a retro dataset covering borrowers whose outcomes are already known, the calibration can be refitted from scratch against the lender's own outcomes. If only the two scores overlap, with no paired outcomes, quantile mapping is the alternative. Where a lender starts overlap-only but expects outcome data later, the two chain: quantile map now, refit when the data supports it.

How does refitting the calibration work?

The calibration layer maps a score into the log-odds space of an account outcome, and its shape is controlled by two parameters, an intercept and a slope. Refitting means estimating both against the lender's own outcomes by logistic regression on the new score, which makes the calibration native to the new score rather than a translation of the old one. This is not exotic work: score-to-outcome recalibration is a routine refresh that lenders already run periodically, and the sensible response to a transition is to run it more often while performance history accumulates.

What is quantile mapping, in one sentence?

A rank-preserving transformation that wraps each borrower's new score onto the old score's scale at the same rank position, so the existing calibration reads a value on the scale it was fitted against. If a borrower's new score sits at the 80th percentile of the new distribution, the calibration is fed the old-score value from the 80th percentile of the old distribution. It needs no outcome data, only a window of borrowers scored under both. Its weakness is its dependence on the reference period chosen: a stale reference sample carries stale patterns into every model that consumes the mapped score.

Does a given score mean the same thing at every lender?

No, and this is the reason a bridge cannot be borrowed from a peer. The same score value implies a different outcome rate on a different book, because the books differ, so each lender's intercept and slope have to reflect its own portfolio. The point has been measured. Butaru, Chen, Clark, Das, Lo and Siddique applied machine learning to account-level credit card data from six major commercial banks covering January 2009 to December 2013 and found that risk factors, their sensitivities and the predictability of delinquency all varied significantly between institutions holding structurally comparable revolving portfolios. Their conclusion was that risk parameters should be calibrated to each institution's own exposures rather than to an industry benchmark.

How should the bridge be monitored once it is live?

Feature by feature on the population crossing it, because the applicant mix moves as the economy and the lender's own marketing change. The conventional metric is the population stability index, which splits the score into bands and compares today's share in each band against a reference sample. The bands are a convention rather than a rule: below 0.10 is read as stability and anything above the upper cut-off as drift material enough to investigate, with that cut-off drawn at 0.20 or 0.25 depending on house practice. Its real limitation is that it is univariate, so it cannot detect a shift in the joint distribution that leaves each feature's own distribution intact, and aggregating several features' indices into one number has no statistical meaning.

Which reference sample should the index compare against?

Both, for different questions. The original build sample shows how far the population has moved since the model was first built, which mixes the effect of the switch with years of ordinary drift. The last sample before the switch isolates the effect of the switch itself. Reporting only the first hides what the migration did; reporting only the second hides how far the model was already from its build population.

How do you tell whether the swap is making the right calls, not just whether the population moved?

Stability monitoring answers the second question and not the first. For the first, run the new model against the old on the same traffic: a shadow engine records the decision each would have made without acting on the new one, and an A/B split directs part of the traffic to each and acts on both. Either produces the two populations worth examining, namely the accounts the old engine accepted and the new one would decline, and the reverse.

Sources

  1. 1 PRA SS1/23. Model risk management principles for banks View source ↗
  2. 2 EBA/RTS/2026/05. material model changes, amending Delegated Regulation (EU) No 529/2014 View source ↗
  3. 3 Butaru, Chen, Clark, Das, Lo and Siddique (2016). Risk and risk management in the credit card industry View source ↗
  4. 4 Djurovic. Representativeness testing and the classifier two-sample test View source ↗
Receive updates directly in your inbox

Stay connected