Data

The role of data in credit modelling

Model performance debates fixate on the algorithm. The bigger gains sit in the data: integrating the sources a lender already has, and engineering features that reflect how customers actually behave.

The hottest topic in credit risk modelling often centres on which algorithm to use. The choice of algorithm matters, but it isn’t where the biggest improvements in performance come from. Logistic and linear regression are valued for their simplicity and transparency. More sophisticated machine learning models, like random forests and gradient boosting machines, hold more predictive power but are harder to interpret. Our guide to machine learning in credit risk modelling weighs that trade-off model by model.

The real progress is in the data

The biggest improvements in credit model performance come from the data itself: how it’s used, and how different sources are brought together to reflect real-world behaviour.

Lenders now have access to far more than traditional application and credit bureau data. Open banking, transactional patterns, behavioural trends, internal customer relationship management (CRM) systems, and even unstructured notes from customer conversations all hold valuable signals.

The next challenge isn’t collecting more data. It’s integrating what we already have.
Where the signal lives: Three tiers of credit data, and the two that are underworked
1

Khandani, Kim and Lo (2010)

Consumer credit-risk models via machine-learning algorithms

View source ↗

The data scientists making the biggest impact aren’t just tuning models. They’re connecting messy datasets, aligning definitions, resolving inconsistencies, and engineering features that reflect how customers actually behave. Often that means aggregating transactions into meaningful summaries, building variables that represent business logic, or cleaning records so that models don’t learn from noise. These steps can drive larger gains than tuning an algorithm. One 2010 study put a number on it: joining a commercial bank’s own customer transaction records to credit bureau data improved delinquency forecasting enough to save an estimated 6 to 25% of total losses, on conservative assumptions about the cost of cutting credit lines.1

The importance of collaboration

One of the most underrated aspects of credit modelling is communication. The best models aren’t built in isolation. They’re made by talking to underwriters, front-line teams, and the people who use them day to day. Those conversations clarify what variables mean in practice, how they’re used in decisions, and what really drives credit risk.

2

Butaru, Chen, Clark, Das, Lo and Siddique (2016)

Risk and risk management in the credit card industry

View source ↗

Without that input, it’s easy to build a model that is technically sound but misaligned with the business. Local knowledge is not a nicety here. One study looked at account-level credit card data from six major commercial banks between January 2009 and December 2013. It found that risk factors, their sensitivities and the predictability of delinquency all varied significantly between institutions holding revolving credit card portfolios.2 A feature set that works at one lender is a hypothesis at the next. It might rely on unstable or unnecessary data, miss important signals, or use variables in ways that make no sense to decision makers. Good data scientists bridge that gap, making sure features are grounded in real-world understanding and that models genuinely support business needs.

What lenders should do next

It’s easy to chase marginal gains through model tuning or the latest algorithm. The less glamorous work is cleaning the data, enriching it, and engineering better features, and it usually carries the larger lift in discriminatory power.

Integrate open banking data

3

BCBS d575

Digitalisation of finance

View source ↗

Traditional bureau data is useful but limited. It gives a slightly delayed view of a customer and often misses key segments, such as thin-file or new borrowers. Open banking data gives a real-time view of income, spending habits, and account activity, and it exists in usable form because regulation made banks expose it. In the UK that was the CMA’s retail banking order and in the EU it was PSD2. The Basel Committee’s May 2024 paper on the digitalisation of finance notes that open banking regimes differ worldwide in whether they are mandatory or voluntary and regulatory led or market driven, and in whether they mandate the use of APIs at all.3 Lenders should prioritise integrating open banking data, and engineer powerful features from it such as income volatility and discretionary spending patterns.

Capture first-party behavioural data

Application forms, call centre notes, CRM data, and product usage logs all provide signals that bureau files can’t. Missed direct debits on other products, the frequency of service interactions, or how quickly applicants respond to document requests can all predict future risk. These sources are often underused, but hold real value when cleaned and engineered properly.

Use AI for conversational data

Call notes, emails, and chat logs often contain early warning signs, such as complaints or income issues, but they’re hard to use at scale. Large language models can now scan and categorise these conversations automatically, flagging risk-relevant events like mentions of financial hardship, broken arrangements, or repeated contact. That turns unstructured communications into structured risk signals a model can learn from, bringing previously untapped data into the credit decision. Reading call notes with a large language model also changes the model’s governance profile.

4

PRA SS1/23

Model risk management principles for banks

View source ↗

The Prudential Regulation Authority (PRA) sets out its model risk expectations in SS1/23. Under Principle 1, a firm’s assessment of model complexity may consider the use of unstructured data alongside interpretability, explainability and transparency, and the potential for designer or data bias.4 Expect a model that reads call notes to sit higher in the firm’s own risk tiering than one that does not.

Strengthen cross-team communication

Involve underwriters, product managers, and credit policy owners early. If a model uses "number of accounts opened in the last three months", ask what it’s really capturing: fraud, opportunism, or normal customer onboarding? That context shapes whether a feature should be included, excluded, or transformed. The best models reflect not just what is predictive, but what is usable in a decision, and the way to get there is a shared understanding between model owners, data scientists, data engineering, operations, and product.

Prioritise data monitoring

As alternative data grows, so does the risk of drift. Lenders need granular, ongoing monitoring of data inputs: tracking changes in open banking connection success rates, for example, or distribution shifts in derived income features.

5

Djurovic

Representativeness testing and the classifier two-sample test

View source ↗

The usual measure of a distribution shift is the population stability index. It compares the share of the population falling in each bucket of a feature today against the share that fell there in a reference period, and returns zero when the two are identical. The industry reads it against conventional thresholds: below 0.10 for stability, 0.10 to 0.25 for modest drift, and above 0.25 for drift material enough to investigate.5 The index works one feature at a time, so it will not catch the case where every feature looks unchanged on its own but the way they move together has shifted.

Build for explainability

6

Lundberg and Lee (2017)

A Unified Approach to Interpreting Model Predictions

View source ↗

Even complex, high-feature-count models need to be explainable. Use tools like SHAP early in development to spot misleading variables, or features that can’t be easily communicated to the business or the regulator. SHAP, short for SHapley Additive exPlanations, attributes a single prediction to the contribution each feature made to it. The method’s standing rests on a 2017 result: among the ways of explaining a prediction as a sum of feature contributions, the Shapley value is the only one that satisfies local accuracy, missingness and consistency together.6

Final thoughts

Predicting credit risk is at the heart of any lending business. Models keep evolving, but the most meaningful improvements now come from getting the data right: pulling in the right sources and creating features that reflect real behaviour, not just what’s easy to model. It also means working closely with the people who use these models every day, so the outputs are trusted, explainable, and aligned with how lending decisions are actually made.

The most successful lenders won’t necessarily be the ones with the most advanced algorithms. They’ll be the ones who know how to turn messy, fragmented data into features a model can learn from and an underwriter can defend.

Frequently asked questions

Where do the biggest gains in credit model performance actually come from?

The data, rather than the algorithm, is where most of the remaining performance sits in a mature credit book. The clearest evidence is Khandani, Kim and Lo's 2010 study in the Journal of Banking and Finance, which combined a major commercial bank's own customer transaction records with credit bureau data covering January 2005 to April 2009. Their forecasts improved the classification of credit card delinquencies and defaults enough that, on conservative assumptions about the cost of cutting credit lines, the authors estimated savings of 6 to 25% of total losses. The lift came from joining two data sources the bank already had, not from a novel algorithm.

Which data does a lender already hold but rarely put into a model?

First-party behavioural data is the most commonly underworked source: customer relationship management records, call centre notes, product usage logs, missed direct debits on other products, and how quickly an applicant returns requested documents. These sit inside the lender rather than at a bureau, which is why they are rarely engineered into features. A missed direct debit on an unrelated product is a live signal about capacity to pay, and no bureau file will carry it for weeks.

What does open banking data add that a credit bureau file cannot?

Open banking data gives a current view of income, spending and account activity, where a bureau file gives a slightly delayed one and carries little at all on new or thin-file borrowers. The Basel Committee's May 2024 paper on the digitalisation of finance defines open banking as the customer-permissioned sharing of banking data with third parties, and notes that regimes differ worldwide in whether they are mandatory or voluntary. In the UK and the EU they are mandatory, through the CMA's retail banking order and PSD2 respectively, which is what shifts the bank's role toward infrastructure provider. The features worth building from it are derived rather than raw, income volatility and discretionary spending patterns being the obvious two.

Can call notes and chat logs be used in a credit model?

Yes, and doing so changes how the model must be governed. Large language models can categorise conversational records at scale, turning mentions of financial hardship, broken arrangements or repeated contact into structured risk signals. The governance consequence is explicit in the PRA's SS1/23: under Principle 1, a firm's assessment of model complexity may consider the use of unstructured data alongside measures of interpretability, explainability and transparency, and the potential for designer or data bias. Bringing conversational data into a model therefore tends to push it up the firm's own risk tiering, which is a documentation and validation cost rather than a reason not to do it.

Why should underwriters be involved in designing a credit model?

Underwriters and front-line teams know what a variable means in the decision, which is knowledge the data alone does not carry. Take "number of accounts opened in the last three months": whether that is capturing fraud, opportunism or ordinary onboarding decides whether the feature belongs in the model, needs transforming, or should be dropped. Without that input a model can be statistically sound and still misaligned with how the business lends, relying on unstable inputs or using variables in ways decision makers cannot act on.

How should a lender monitor alternative data for drift?

Monitoring should run feature by feature on the inputs, not only on the model's output. The conventional metric is the population stability index, and the industry reads it against fixed bands: below 0.10 indicates stability, 0.10 to 0.25 modest drift, and above 0.25 material drift that needs investigating. Its limitation is worth knowing before relying on it. The index is univariate, so it cannot detect a shift in the joint distribution that leaves each feature's marginal distribution intact, and aggregating several features' indices into one number has no statistical meaning. For alternative data specifically, the operational inputs matter as much as the statistical ones: a fall in open banking connection success rates changes the population a model scores.

What makes a high-feature-count model explainable enough for a regulator?

Attribution at the level of the individual decision is what turns a complex model into an explainable one. Lundberg and Lee's 2017 paper established the theoretical basis: among methods that explain a prediction as a sum of feature contributions, the Shapley value is the only one that satisfies local accuracy, missingness and consistency at the same time. The practical consequence is that a credit officer can be shown which features pushed a given application's predicted probability of default up and which pulled it down, and by how much, in the model's own output units. Later work by the same group gave tree ensembles an exact algorithm fast enough to run on a large boosted model, which is what made the method practical there. No supervisor has mandated the method. What creates the incentive is the PRA's model risk management principles, which require complexity to be assessed and governed, and the FCA's Consumer Duty, which expects analytics-driven decisions to be communicated to customers in a way they can understand.

Does a credit model that performs well at one lender transfer to another?

Not reliably, even across portfolios that look the same. Butaru, Chen, Clark, Das, Lo and Siddique applied machine learning to account-level credit card data from six major commercial banks covering January 2009 to December 2013, and found that risk factors, their sensitivities and the predictability of delinquency all varied significantly between institutions holding the same type of revolving credit card exposures. Their conclusion was that capital ratios, loss reserves and risk parameters should be calibrated to each institution's own exposures and forecasts rather than to an industry benchmark. The practical reading for a lender is that a competitor's feature set is a hypothesis to test on your own data, not a specification to copy.

Sources

  1. 1 Khandani, Kim and Lo (2010). Consumer credit-risk models via machine-learning algorithms View source ↗
  2. 2 Butaru, Chen, Clark, Das, Lo and Siddique (2016). Risk and risk management in the credit card industry View source ↗
  3. 3 BCBS d575. Digitalisation of finance View source ↗
  4. 4 PRA SS1/23. Model risk management principles for banks View source ↗
  5. 5 Djurovic. Representativeness testing and the classifier two-sample test View source ↗
  6. 6 Lundberg and Lee (2017). A Unified Approach to Interpreting Model Predictions View source ↗
Receive updates directly in your inbox

Stay connected