AI use case

AI credit scoring with alternative data for thin file applicants

A machine learning credit model that adds consumer permissioned alternative data, such as bank account cash flow, rent, utility and telco payments or ecosystem data, to credit bureau data, so a lender can assess applicants with thin or no credit files and return a decision with specific reasons.

By Len Debets · Last verified 26 September 2026 · 5 public deployments

USD 450,000 to USD 4.8 million
Indicative value per year
A consumer lender receiving 200,000 personal loan applications a year. Worked example, see how it is calculated.

What problem does it solve?

A bureau score needs a credit history. Young adults, migrants, gig workers, the self employed and people who simply never borrowed have little or none, so a traditional scorecard either declines them or prices them as if they were high risk. In the United States alone, the CFPB estimates that 26 million adults have no credit history at a nationwide credit bureau and another 19 million have one too thin or stale to score.

The information that would show whether these people can repay usually exists: salary and expenses in their bank account, years of rent, utility and phone payments, or activity on a super app. Lenders could not use it at scale because it was unstructured, scattered and not permissioned for credit. Open banking, consent frameworks and machine learning now make it usable, but they also bring new questions about fairness, explainability and privacy that a bureau scorecard never raised.

How does it work?

  1. Ask for consent. The applicant chooses to share additional data, such as a bank account connection through open banking or data from an ecosystem partner, and is told what it is used for.
  2. Turn raw data into features. Transactions are categorised into income, rent, essential spend, debt payments and overdraft use; stability and trend features are computed over several months.
  3. Score alongside the bureau. A machine learning model, or a cash flow score added to the existing scorecard, estimates default risk using both bureau and alternative features.
  4. Decide within policy. A decision engine applies credit policy, affordability rules and limits, approves, declines or refers the case, and sets line size and price.
  5. Explain the outcome. Each decline or unfavourable term carries the specific principal reasons derived from the model, and the applicant can ask what would change the outcome.
  6. Monitor. Approval rates, default rates and outcomes by protected group are tracked per segment, and the model is revalidated when data or populations drift.
Audience
Back office
Autonomy
Supervised agent
Adoption
Early adopters
Channels
API and system to system, Mobile app, Web chat

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Inclusion and access, Revenue growth, Risk and loss reduction, Speed and cycle time.

Indicative value

A consumer lender receiving 200,000 personal loan applications a year

USD 450,000 to USD 4.8 million

Net contribution from additional approvals per year

How this is calculated

Formula: applications * declineShare * rescuedShare * contributionPerLoan. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Applications per year applications, applications per year200,000200,000The reference lender.
Share of applications declined today declineShare, fraction of applications0.30.4Editorial assumption for a mainstream personal loan book. Replace with your own decline rate.
Share of declines that alternative data turns into sound approvals rescuedShare, fraction of declines0.050.15The high end is the Atlanticus figure, where the vendor reports that 15% of marginal declines (not all declines) could be approved profitably, so it is an upper bound; the low end allows for weaker data coverage and consent drop off. Editorial assumption, replace with your own.
Net contribution per additional approved loan over its life contributionPerLoan, USD per loan150400Editorial assumption after expected credit losses and funding cost. Replace with your own.

What it leaves out: Leaves out the cost of data access, model development and validation, the consent drop off rate, and any change in losses on loans the bank would have approved anyway. The uplift must be proven on your own population with a holdout before it is counted.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Upstart Network

United States · Banking · 2019

ScaledGrade B

Upstart underwrites and prices consumer loans with a machine learning model that adds alternative data, such as education and employment, to traditional credit data. In 2017 it received the CFPB's first No Action Letter for this model, and in 2019 the CFPB published the access to credit results Upstart reported under that letter: in simulations run against a hypothetical traditional model on the same applicant pool, Upstart's model approved 27% more applicants with 16% lower average APRs, across all tested race, ethnicity and sex segments. The CFPB states it did not separately replicate these simulations. An independent fair lending monitorship of the model published its initial report in April 2021.

No outcome disclosed.

Atlanticus

United States · Banking · 2025

ProductionGrade C

Atlanticus, which runs credit card brands through bank partners for consumers overlooked by prime lenders, added consumer permissioned bank transaction data (Nova Credit Cash Atlas) to its decisions, with consent collected inside an embedded finance marketplace. The vendor reports that 15% of marginal declines, applicants that bureau data alone would have rejected, could be approved profitably with cash flow insights, with no deterioration in credit quality, and that the data is also used to set line sizes and pricing.

No outcome disclosed.

Patelco Credit Union

United States · Banking · 2025

PilotGrade C

Patelco Credit Union tested VantageScore 4plus, a score that combines credit file data with consumer permissioned open banking (bank account cash flow) data, on its own portfolio. In the pilot, 12% of subprime and 15% of near prime members moved to higher credit tiers, and predictive power in originations improved by 4.8% over VantageScore 3.0. It was a portfolio test, not a production rollout. The larger figures in the press release headline (33% and 41%) come from the second pilot at Michigan State University Federal Credit Union, not from Patelco.

No outcome disclosed.

Golden 1 Credit Union

United States · Banking · 2024

ProductionGrade C

Golden 1, a California credit union with about USD 21 billion in assets, built a custom machine learning credit scorecard with Zest AI, trained on its own members and on other Californians who resembled its membership. It launched on credit cards in December 2022 and extended to unsecured and auto loans. The CEO reported higher approvals overall and a 28% increase in approvals to protected classes of borrowers. The article does not say that the scorecard uses data from outside the credit file, so the record illustrates the machine learning half of this use case.

No outcome disclosed.

GXS Bank

Singapore · Banking · 2024

ProductionGrade C

GXS Bank, a Singapore digital bank, decisions its FlexiLoan personal loan with user permissioned data from its ecosystem partners Grab and Singtel layered on top of credit bureau scores, to expand credit access to underserved users, such as people starting their careers and entrepreneurs with fluctuating incomes, who were previously overlooked by traditional banks. The FICO decision platform returns credit decisions in milliseconds, and the bank reports onboarding in under three minutes for the vast majority of approved applications. The platform was implemented in three months.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Historical applications with outcomes (performance over at least 12 months) for model training and validation
  • Access to permissioned alternative data with a clear legal basis, such as open banking or partner data
  • Protected attribute data or accepted proxies for fairness testing, where the law allows
  • Documented credit policy and affordability rules

Systems to integrate

  • Open banking or account aggregation provider
  • Credit bureau
  • Decision engine or loan origination system
  • Model monitoring and model inventory
  • Customer channels for consent and status updates

Complexity: High

The model is the easy part. The work is in consent flows that people complete, data coverage, model validation, fair lending testing, reason codes that stay specific and accurate, and integration with the existing decision engine and bureau.

  1. 1

    Pick the segment and the question

    Start with one product and one segment, such as thin file applicants for a personal loan, and one question: can alternative data safely approve some of today's declines?

  2. 2

    Backtest before you lend

    Score past applicants with and without the new data and compare default rates at equal approval rates. The published VantageScore pilots at Patelco Credit Union and Michigan State University Federal Credit Union tested the score on existing portfolios before lending on it, and that is the right first step.

  3. 3

    Design consent as a product

    The uplift only reaches customers who share their data. Explain the benefit in plain words, keep the connection step short and offer it at the point of decline or referral.

  4. 4

    Build reasons in from the start

    Choose features and methods that let you state the specific principal reasons for every adverse action. Test the reasons with real applicants before launch.

  5. 5

    Test for fairness and validate independently

    Run disparate impact analysis and search for less discriminatory alternatives, then put the model through the same independent validation as any credit model.

  6. 6

    Launch with a champion and challenger

    Route a share of traffic to the new model, keep the old one as control, and widen only when loss rates on the new approvals are confirmed.

Guardrails

  • Alternative data only with explicit, recorded consent and a documented legal basis
  • No feature that acts as a proxy for a protected characteristic, checked by testing, not by assertion
  • Specific, accurate reasons for every decline or unfavourable change in terms
  • Credit policy and affordability rules stay deterministic and outside the model
  • A human owns referred cases and any appeal against a decision

KPIs to instrument

  • Approval rate on the target segment versus a control group
  • Default and loss rates of the incremental approvals over time
  • Consent completion rate at the data sharing step
  • Approval and pricing gaps across protected groups
  • Share of decisions returned automatically, and decision time

Human in the loop

Credit officers decide referred and borderline cases and handle appeals. Model risk and fair lending teams approve the model before launch and review performance by segment every quarter. The customer can ask for human review of an automated decision.

Common failure modes

Uplift that disappears in production
A backtest on past applicants does not match who actually consents. Measure with a live control group, not only on history.
Proxy discrimination
Behavioural or device data can stand in for race, sex or age. Test outcomes by group and remove features that drive unjustified gaps.
Reasons nobody can act on
Generic reasons such as "failed to achieve a qualifying score" or "based on internal standards or policies" do not meet the requirement for specific reasons and frustrate applicants. Map features to plain, specific reasons.
Data that goes stale or disappears
A partner or aggregator changes coverage and the model degrades silently. Monitor input distributions and fall back to the bureau model.

What are the risks and rules?

EU AI Act

High risk

Annex III point 5(b): AI systems intended to evaluate the creditworthiness of natural persons or establish their credit score are high risk, except systems used to detect financial fraud. Providers need risk management, data governance, logging and human oversight. Deployers must carry out a fundamental rights impact assessment before use (Article 27), and affected persons have a right to an explanation of individual decisions from the deployer (Article 86).

Guidance

Controls to put in place

  • Model inventory entry with an accountable owner, validation report and approved use
  • Consent records linked to every decision that used alternative data
  • Fair lending testing before launch and on a fixed schedule, with a documented search for less discriminatory alternatives
  • Reason code library reviewed by compliance and tested for accuracy against model output
  • Drift and performance monitoring with a documented fallback to the bureau model

When it went wrong elsewhere

Frequently asked questions

How much does alternative data increase approvals?
It depends on the population and the data. In simulations it reported to the CFPB, which the CFPB did not separately replicate, Upstart's model approved 27% more applicants than a hypothetical traditional model, with lower average APRs. In a 2025 open banking pilot, Patelco Credit Union saw 12% of subprime and 15% of near prime members move to a higher credit tier. Prove it on your own applicants with a control group.
Is alternative data credit scoring high risk under the EU AI Act?
Yes, when it evaluates the creditworthiness of natural persons (Annex III point 5(b)). That brings obligations for risk management, data governance, human oversight and logging, deployers must assess the impact on fundamental rights, and affected people have a right to an explanation of the decision.
Does using machine learning excuse a lender from giving specific decline reasons?
No. Regulation B requires a statement of reasons that is specific and indicates the principal reasons for the adverse action; saying the applicant failed to reach a qualifying score or cites the creditor's internal standards is not enough. Design the model and the reason codes together.
Which alternative data is safest to start with?
Consumer permissioned bank transaction data is a common starting point, because it measures income and spending directly and the applicant chooses to share it; the Atlanticus and Patelco examples on this page both use it. Device, social and behavioural data carry higher privacy and proxy discrimination risk.

How to cite this page

Blits.ai AI Use Case Library, "AI credit scoring with alternative data for thin file applicants", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/alternative-data-credit-scoring. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Banking

AI cash flow underwriting for small business loans

An underwriting engine that assesses a small business's repayment capacity from live bank transactions, point of sale and payment flows, receivables and accounting data instead of audited accounts, and returns a decision recommendation with the evidence and reasons behind it.

Deployments
4 public, best grade B
Autonomy
Supervised agent
BankingPayments and cards

AI drafted explanations for credit declines and adverse actions

An assistant that turns the reason codes of a credit model into an accurate, specific and readable explanation of a decline, reduced limit or repricing for the customer, and a matching internal rationale for the file, without adding any reason the model did not produce.

Deployments
2 public, best grade B
Autonomy
Copilot
BankingPayments and cards

AI for application and identity fraud detection

AI that checks incoming account and loan applications for forged or AI generated documents, synthetic and stolen identities, and coordinated application rings, by analysing documents, device and application data across the whole queue and cross checking against bureau and official sources.

Deployments
6 public, best grade B
Reported detection improvement
2.5x
Department for Work and Pensions, organization claim
Banking

Conversational AI for loan application intake

A conversational assistant on web, app, messaging or voice that explains loan products, captures the application through dialogue in the customer's language, checks documents and basic eligibility rules, and hands a complete, structured application to origination, without making the credit decision.

Deployments
5 public, best grade C
Autonomy
Supervised agent
BankingInsurance

AI copilot for model risk validation and monitoring

A copilot for independent model validation and review, whether run by a bank's validation function, an external tester or a supervisor, that checks model documentation against the model risk standard, generates and scores challenger tests (for generative AI, often with an LLM as a judge calibrated against human experts), watches production models for drift and drafts and consistency checks the validation report. An accountable validator owns every conclusion.

Deployments
3 public, best grade B
Autonomy
Copilot