Some sources create more risk because they are easy to manipulate, unevenly predictive, or legally inappropriate for the credit context. The article points to social media as a poor fit for consumer lending because it can expose protected characteristics and is hard to dispute. By contrast, recurring payment and account activity are more closely tied to financial behavior and default risk.
Why certain alternative data sources are riskier in credit underwriting
Some alternative data creates more risk because the signal is weak, the data can be gamed, and the source may be inappropriate for a lending decision. Social, behavioral, or other indirectly related data can look predictive in a model without being robust, fair, or defensible once it is challenged by adverse selection, manipulation, or compliance review.
That makes the core issue not whether a data source is “new,” but whether it can reliably support a credit conclusion without introducing hidden bias, weak explainability, or legal exposure. Better sources tend to reflect payment behavior, cash flow, or account activity that is more directly tied to repayment capacity.
What makes a data source low-trust in practice
Low-trust sources often fail on one or more practical tests. They may be easy for consumers to alter, noisy enough that the model learns correlation instead of credit behavior, or sensitive enough that their use raises concern even if they improve a score in backtesting. If a borrower can change the input cheaply, the data is vulnerable to strategic behavior. If the lender cannot explain why it matters, the source becomes harder to defend in reviews and disputes.
Another issue is that not all useful signals belong in all lending contexts. A source can be interesting from a data-science perspective but still be a poor fit for underwriting because it introduces protected-class risk, weakens adverse action explainability, or creates a mismatch between what is being measured and what credit risk actually means.
EU General Data Protection Regulation (GDPR) is relevant here because data minimization, purpose limitation, and fairness constraints shape how far lenders can go when repurposing personal data for credit decisions.
Why payment and account data is usually safer than social data
Recurring payment, deposit, and account activity usually create less risk because they are closer to the underlying economic question: will this person repay? These signals are not perfect, but they are easier to justify, easier to test for stability, and more naturally connected to default behavior than lifestyle or social presence.
By contrast, social media and similar sources often fail on relevance and defensibility. They can contain proxies for protected characteristics, they can be sparse or performative, and they may not generalize across populations or time periods. A lender can sometimes use them in research or fraud analytics, but using them as a core underwriting signal increases the chance that the model reflects noise, demographic bias, or behavior that has little durable relationship to credit performance.
NIST Privacy Framework helps frame this trade-off by emphasizing governance over collection, use, and downstream privacy risk, while NIST SP 800-53 Rev 5 Security and Privacy Controls supports the broader control expectation around privacy, auditability, and controlled use of sensitive information.
Why underwriting teams should separate prediction from defensibility
A source can improve model performance and still be the wrong choice for production lending. The underwriting question is not only “does it predict?” but also “can we justify it, monitor it, and defend it when challenged?” That means the team should look for stable performance across segments, clear linkage to repayment behavior, low susceptibility to manipulation, and a clean story for compliance and adverse action.
Current guidance suggests that the best signals are those that are both behaviorally relevant and operationally stable. If a source is highly predictive only in one slice of the portfolio, or only when combined with opaque transformations, it may be too fragile for a credit policy even if it looks attractive in a pilot.
NIST Privacy Framework also supports the governance side of this decision, because defensible underwriting depends on data practices that can be explained and bounded, not just on model accuracy.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the technical controls, while GDPR defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| GDPR | A.5.15 — Data processing and privacy | Credit use of personal data must respect purpose and fairness limits. |
| Recommendation — Limit underwriting data use to lawful, necessary, and explainable purposes. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Record Review, Analysis, and Reporting | Defensible credit decisions need reviewable evidence for challenged inputs. |
| Recommendation — Retain decision evidence that shows how data influenced the score. | ||
| NIST CSF 2.0 | GV.OV-01 — Oversight of risk and security | Credit model inputs require governance over risk, fairness, and accountability. |
| Recommendation — Establish oversight for alternative data approval and monitoring. | ||
Practitioner Guidance
What to verify: Before approving a new source, check whether it is directly tied to repayment capacity, whether a borrower can manipulate it cheaply, and whether the same signal can be obtained from a less sensitive source. If the answer relies on proxies for lifestyle or identity rather than observed payment behavior, treat it as high-friction and likely unsuitable for core underwriting.
Decision rule: If a source is strong for prediction but weak for explanation, fairness, or challenge response, keep it out of the lending decision path or restrict it to exploratory analytics. If it is both predictive and clearly linked to financial behavior, it is a much better candidate for credit use.
Practitioner takeaway: The safest alternative data for credit is usually the data that is hardest to game, easiest to explain, and most directly connected to repayment, not the data that is merely easiest to collect.
Related resources from NHI Mgmt Group
- Why do data sources still create secrets risk in Terraform?
- Why do alternative data models create more compliance risk than traditional scorecards?
- Why do AI agents create new risk when they interact with Zapier-connected data sources?
- Why does stored credit card data in CRM systems create compliance and breach risk?