Join our Newsletter — 33% off our NHI Course

How should lenders decide which alternative data sources belong in a creditworthiness model?

Lenders should choose alternative data only when it improves prediction, complies with credit rules, and matches the decision they are trying to make. Hard data such as payment behavior is usually more reliable than soft data like social signals. The best sources are accurate, timely, specific, broadly covered, and additive to bureau data. If a source does not improve the signal, it should be excluded.

What makes an alternative data source worth using in a credit model?

The deciding question is not whether the data is novel, but whether it improves underwriting in a way that is measurable and defensible. Lenders need a source that adds signal beyond bureau data, fits the decision being made, and can be explained if challenged by compliance, risk, or model governance teams.

That means separating useful predictive features from data that is merely available. A source can be operationally interesting and still fail the model test if it is noisy, hard to verify, or too weakly tied to repayment behaviour to justify use.

How to judge signal quality, coverage, and decision fit

The first filter is predictive value. If the alternative source does not improve out-of-sample performance, calibration, or stability, it is not contributing enough to justify collection and use. Lenders should also check whether the source measures something close to the credit decision itself, rather than a distant proxy that only appears useful in historical data.

Coverage matters just as much as lift. A source that works only for a narrow subgroup, sparse segment, or short time window can create brittle models that perform well in testing but fail in production. Timeliness and granularity matter too, because stale or overly coarse data often lags the borrower’s actual behaviour and can distort risk scoring.

Accuracy and consistency are the other practical test. If the source contains missing values, frequent revisions, weak lineage, or inconsistent definitions across providers, it can degrade the model even when it looks predictive at first glance. In credit decisions, the strongest alternative sources are usually those with clear provenance, repeatable collection, and a direct relationship to observed repayment or affordability.

Which sources usually belong, and which ones usually do not

Hard behavioural data generally belongs earlier in the evaluation queue than soft or inferred signals. Payment history, cash-flow patterns, transaction behaviour, and verified account activity are easier to defend because they are closer to financial capacity and obligation management. By contrast, social signals, broad lifestyle inferences, and loosely correlated digital traces may add little genuine credit insight and can introduce fairness, explainability, and governance problems.

The practical issue is not whether a source is traditional or alternative. It is whether the data is specific enough to the borrower’s repayment ability, reliable enough to support repeatable decisions, and stable enough to survive portfolio monitoring. If the answer is no, the source should stay out of the model even if it is easy to obtain.

For lenders working in regulated environments, the decision also has to align with permissible-use rules, adverse-action readiness, and internal model documentation. The more indirect the signal, the more carefully the lender should justify why it is needed and how it will be reviewed for drift, bias, or unintended proxy effects.

Risk and Threat Considerations

Alternative data creates model risk when the source is weakly linked to repayment behaviour, and it creates governance risk when teams cannot explain why the feature belongs in a lending decision. In practice, the largest failure modes are overfitting, proxy discrimination, stale inputs, and dependence on data that can disappear or change without warning.

Failure mechanism: The model appears to improve in testing because the data captures noisy correlates or short-lived patterns, but those patterns do not hold in production or across borrower populations. If the source is also poorly governed, a lender may not detect drift, coverage gaps, or compliance issues until after decisions have already been made.

Impact: Borrowers can be scored inaccurately, some segments can be unfairly advantaged or excluded, and the institution can inherit explainability, compliance, and reputational exposure. The larger the portfolio, the more expensive these errors become to unwind.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST SP 800-53 Rev 5 sets the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.

Framework Control / Reference Relevance
NIST SP 800-53 Rev 5 CA-7 — Continuous Monitoring Alternative data must be monitored for drift and ongoing model performance.
RA-3 — Risk Assessment Lenders need structured assessment of model, compliance, and fairness risk from new data sources.
PM-31 — Continuous Monitoring Strategy Ongoing oversight is needed when alternative data feeds affect lending decisions.
Recommendation — Monitor alternative data features for drift, coverage loss, and performance degradation over time. Assess each alternative data source for predictive value, bias, and governance risk before use. Define monitoring expectations for alternative data quality, drift, and decision impact.
ISO/IEC 27001:2022 A.5.31 — Legal, statutory, regulatory and contractual requirements Credit models using alternative data must align with applicable lending and data-use rules.
A.5.12 — Classification of information Source data should be classified by sensitivity and business impact before model inclusion.
Recommendation — Map each data source to the legal and contractual limits that govern its use in lending. Classify candidate data sources so sensitive or high-impact inputs receive stricter controls.

Practitioner Guidance

What to verify: Require evidence that each candidate source improves a holdout test, remains stable across time and segments, and can be mapped to a defensible credit rationale. If you cannot explain the mechanism by which the data improves underwriting, it is probably not model-worthy.

Decision rule: Prefer sources that are both predictive and operationally governable, then reject any feature that only adds marginal lift while increasing compliance, fairness, or explainability burden. When two sources perform similarly, choose the one with cleaner lineage, better update cadence, and clearer relevance to repayment behaviour.

Practitioner takeaway: The best alternative data is not the most novel data, it is the data that improves decision quality without weakening the lender’s ability to defend, monitor, and correct the model.