AI use case

AI risk scoring in child welfare intake and investigations

A predictive model that scores a family's risk in the child welfare system, using case history and administrative data, so that staff see a consistent, data informed signal alongside their own judgment, whether that is a call screener and supervisor deciding at referral whether to open an investigation, or a supervisor and social worker deciding how to respond early in one that is already open. The model informs the decision; it does not make it.

By Len Debets · Last verified 29 September 2026 · 2 public deployments

USD 20,000 to USD 202,500
Indicative value per year
A child welfare agency handling 40,000 referrals a year. Worked example, see how it is calculated.

What problem does it solve?

Child welfare agencies receive a very large number of reports on a hotline, online or in person. A call screener has to decide whether to open an investigation, and once an investigation is open, a supervisor and social worker have only a short window to assess safety and plan a response. Los Angeles County's hotline alone received 168,045 calls in 2021, and about 750 emergency response social workers completed 43,505 investigations involving 86,487 children. Neither a missed risk nor an unnecessary investigation is a small thing: one can leave a child unprotected, the other can put a family through a distressing and stigmatizing process for nothing.

Predictive risk models built from case history try to give staff a consistent, data informed signal alongside the referral or the open case, whether at the point a screener decides to open an investigation or in the early days of one that is already open. The same idea has drawn serious scrutiny: a model trained on historic investigation and removal decisions can encode whatever bias sat inside those decisions, and a family cannot see or challenge a score the way they can challenge a person's stated reasoning. Allegheny County, whose tool scores referrals at intake, commissioned independent process and impact evaluations as part of the tool's rollout and published its methodology. Los Angeles County, whose tool supports investigations that are already open, published its own methodology report, implementation insights and quantitative data for investigations, and built a Racial Feedback Equity Loop to check for disproportionate impact on African American families.

How does it work?

  1. A referral arrives. A report of suspected abuse or neglect comes in by phone, online or in person to the agency's intake or hotline service.
  2. The system pulls linked history. Using an integrated data warehouse, the model draws in the family's and any named adults' history across child welfare and, where lawfully available, other public systems.
  3. The model scores the case. A risk score estimates the likelihood of a defined future harm or need, based on patterns in that linked history.
  4. Staff see the score alongside the case. Depending on the deployment, that is a call screener and supervisor deciding whether to open an investigation, or a supervisor and social worker already working an open one. The score is presented alongside, never instead of, the referral or case file and the worker's own read of it.
  5. A person decides. The screener and supervisor decide whether to open a formal investigation, offer services or referrals, or take no further action; or, once an investigation is open, the supervisor and social worker decide how to respond and plan services. Either way, the reasoning is recorded.
  6. The decision and the score are logged and reviewed. Outcomes are tracked over time, including by protected characteristics, to check the model and the process for disparate impact.
Audience
Employee facing
Autonomy
Assist
Adoption
Emerging
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Risk and loss reduction, Employee productivity.

Indicative value

A child welfare agency handling 40,000 referrals a year

USD 20,000 to USD 202,500

Screener time cost avoided on referral triage per year

How this is calculated

Formula: referrals * (screeningMinutesPerReferral / 60) * timeSavedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Referrals screened per year referrals, referrals per year40,00040,000Editorial assumption for a large county agency, replace with your own referral or call volume. Not anchored to a specific organization on this page, because Allegheny County and Los Angeles County publish figures at different stages (calls, referrals and investigations) that are not directly comparable.
Screener and supervisor minutes per referral screeningMinutesPerReferral, minutes per referral2045Editorial assumption for the time to review a referral with its history. Replace with your own time and motion data.
Share of screening time saved by having the score and linked history ready timeSavedShare, fraction of screening time0.050.15Conservative editorial assumption. Allegheny County's published impact evaluation measured effects on screening and case opening decisions, not on screening time, and neither organization on this page has published a measured screening time saving. Replace with your own measurement.
Cost of a screener hour, fully loaded costPerHour, USD per hour3045Editorial assumption for a fully loaded child welfare caseworker cost. Replace with your own.

What it leaves out: This values only time saved gathering and reviewing information for the decision. It leaves out the cost of building, validating and continuously monitoring the model for bias. Allegheny County commissioned an independent impact evaluation of its tool (Jeremy Goldhaber-Fiebert, Stanford University, published by Allegheny County DHS, April 2019), which found the AFST increased the accuracy of decisions to advance children to investigation and reduced disparities in case opening rates between Black and white children, but found "no evidence that the AFST resulted in greater screening consistency" within individual call screeners. Los Angeles County has not published a comparable before and after evaluation of its own tool's effect on decisions.

Who already uses it?

2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Los Angeles County Department of Children and Family Services

United States · Government and public sector · 2021

PilotGrade B

Los Angeles County's Department of Children and Family Services (DCFS) launched a risk stratification pilot in August 2021 in three regional offices (Belvedere, Lancaster and Santa Fe Springs). The model was developed following an analysis conducted by DCFS and university based researchers with the Children's Data Network, who established that a relatively small number of investigations show chronic patterns of alleged abuse or neglect that signal significant service needs for families. The tool delivers information to supervisors at the outset of an investigation that is already open, so they and social workers have it early, when they have only a short window to assess safety and plan services; it does not score incoming hotline calls or help decide whether to open an investigation. DCFS built a "Racial Feedback Equity Loop" into the model to identify screening practices and community reporting patterns that may result in unnecessary investigations disproportionately burdening African American families, and published a methodology report, implementation insights, quantitative data for investigations and an ethical review of the tool's use case rather than only an internal report. As of the source date, August 2022, DCFS had not decided whether to expand the model beyond the pilot offices.

No outcome disclosed.

Allegheny County Department of Human Services

United States · Government and public sector · 2016

ScaledGrade B

Allegheny County's Department of Human Services (DHS) has used the Allegheny Family Screening Tool (AFST) in its child welfare intake office since August 2016 to support call screeners and their supervisors when a report of suspected child abuse or neglect comes in. The tool generates a risk score from data in the county's own data warehouse; staff use the score, alongside their own experience and clinical judgment, to decide whether to open a formal investigation, offer services or take no further action. DHS commissioned independent process, impact and ethical evaluations before and after go live and has since advised other jurisdictions considering similar tools. The 2019 impact evaluation, led by Jeremy Goldhaber-Fiebert of Stanford University, found the AFST increased the accuracy of decisions to advance children to investigation and reduced disparities in case opening rates between Black and white children, but found "no evidence that the AFST resulted in greater screening consistency" between call screeners.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Multiple years of linked case history across child welfare, ideally in an integrated data warehouse
  • A validated method for joining records to the same family or individual reliably
  • An independent validation and ethical review completed before any live use

Systems to integrate

  • Case management system used by call screeners, investigators and supervisors
  • The agency's data warehouse or integration layer linking welfare and related records
  • Fairness and outcome monitoring reporting for the model's owner

Complexity: High

The model needs years of linked, high quality administrative data across child welfare and, where lawful, related systems, a validation and fairness evaluation before it ever sees a live referral, and an ongoing monitoring program, because an error in this domain can mean a missed protection concern on one side or an unnecessary investigation of a family on the other.

  1. 1

    Commission an independent evaluation before any live score

    Have the model, its data and its likely disparate impact validated by an outside evaluator on historic data, and publish the methodology, before a screener ever sees a live score.

  2. 2

    Present the score next to the referral, never instead of it

    Show screeners the score and, where possible, the reasons behind it, alongside the full referral and case history, so it is one input among several, not a replacement for reading the case.

  3. 3

    Write a decision rule that keeps a person accountable

    Require a screener and a supervisor to make and record the screen in or screen out decision and their reasoning, and make it normal, not exceptional, for staff to depart from the score.

  4. 4

    Build a fairness feedback loop from day one

    Track outcomes by race and other protected characteristics, and act on what they show rather than only reporting them. Los Angeles County's Racial Feedback Equity Loop uses data this way to identify screening practices and community reporting patterns that may result in unnecessary investigations disproportionately burdening African American families.

  5. 5

    Publish what you can

    Release what you can. Allegheny County commissioned independent process and impact evaluations and published its methodology; Los Angeles County published its methodology report, implementation insights and quantitative data for investigations. Either practice lets families, advocates and researchers hold the program to account.

  6. 6

    Treat a pilot as a pilot before scaling

    Gather at least a year of paired decisions and outcomes, in a limited number of offices, before deciding whether to expand, and set out in advance what would justify not expanding.

Guardrails

  • A screener and a supervisor make and record every screen in or screen out decision; the score informs it, it does not make it
  • Independent validation and an ethical review completed before the model sees a live referral
  • Regular fairness monitoring by race and other protected characteristics, feeding back into the model or the process
  • Published methodology and, where lawful, anonymized outcome data for external scrutiny

KPIs to instrument

  • Screen in and screen out rates against the model's score, including how often and why staff depart from it
  • Repeat referral or subsequent investigation rates for families who were screened out
  • Fairness metrics across race and other protected characteristics, tracked over time
  • Time from referral to decision

Human in the loop

A call screener and a supervisor decide whether to open an investigation, offer services, or take no further action, using the score alongside the full referral, the family's case history and their own professional judgment; where a tool operates inside an already open investigation, a supervisor and social worker use the score the same way to plan the response. Allegheny County describes its tool as designed "to support, not replace, professional judgment," and its own impact evaluation found that for referrals scoring in the AFST's "mandatory" screen in range, which is meant to prompt a screen in decision, only 61 percent were in fact screened in between December 2016 and November 2018, evidence that screeners kept and used their discretion even at the top of the score range.

Common failure modes

The score becomes the decision in practice
Under caseload pressure, screeners start following the score without engaging with the underlying referral, so human oversight exists on paper only. Audit a sample of decisions against the full case, not just the score used.
Historic bias in the training data becomes bias in the score
A model trained on past investigation or removal decisions can reproduce whatever bias sat inside those decisions, at scale. Validate and monitor by protected characteristic before launch and continuously after it.
A population level signal gets read as a finding about one family
The score estimates statistical risk across similar cases; it is not evidence about what is actually happening in this family. Train staff explicitly on that distinction and require them to write their own reasoning, not just cite the score.
No plan for when to stop
Predictive risk tools in this field draw close public and legal scrutiny. Decide in advance what evidence, such as an adverse fairness finding or an independent evaluation result, would trigger pausing or ending the tool, and who has the authority to do it.

What are the risks and rules?

EU AI Act

High risk

Annex III point 5(a): systems used by or on behalf of a public authority to evaluate a natural person's eligibility for essential public assistance benefits and services, or to grant, reduce, revoke or reclaim them. Recital 58 names "social services providing protection" among those essential services, and both tools sit inside exactly that: whether a family receives a child protective investigation and the services that can follow. Article 6(3) lets an Annex III system avoid the high risk tier when it only performs a narrow procedural task, improves a completed human decision, or is a preparatory step, but that exemption does not apply where the system performs profiling of natural persons, which this model does from personal and administrative data. So profiling closes the exemption; it is not a separate route into the high risk tier on its own.

Controls to put in place

  • Independent validation and periodic revalidation of the model, published where possible
  • A documented decision rule requiring a screener and a supervisor, with the score as one input among several
  • Regular fairness and outcome monitoring across protected characteristics, with a defined escalation and change process
  • A public, published methodology and, where lawful, anonymized outcome data

Frequently asked questions

Does the AI decide whether to investigate a family?
No. At Allegheny County, a call screener and a supervisor decide whether to open an investigation, using the score alongside the full referral and case history; Allegheny describes the tool as designed "to support, not replace, professional judgment." At Los Angeles County, the score reaches supervisors and social workers early in an investigation that is already open, to help shape the response, not a decision on whether to open it.
How is bias in the risk score managed?
Allegheny County commissioned independent process and impact evaluations before and after launch. Los Angeles County built a Racial Feedback Equity Loop to identify screening practices and community reporting patterns that may result in unnecessary investigations disproportionately burdening African American families. Both agencies published their methodology rather than keeping it internal.
Which agencies have deployed this, and how far?
Allegheny County, Pennsylvania has used the Allegheny Family Screening Tool in its child welfare intake office since August 2016. Los Angeles County piloted a separate risk stratification model in three regional offices starting August 2021; as of its August 2022 update, the county had not decided whether to expand it.
Is this high risk under the EU AI Act?
Yes, under Annex III point 5(a): a system used by a public authority to evaluate eligibility for essential public assistance benefits and services, or to grant, reduce, revoke or reclaim them, which covers whether a family receives a child protective investigation and response. Because this model also profiles families from personal and administrative data, it cannot claim the Article 6(3) exemption some Annex III systems get for narrow procedural or preparatory tasks. Either way it brings requirements for risk management, data governance, logging, human oversight and conformity assessment, even though the final decision stays with a person.

How to cite this page

Blits.ai AI Use Case Library, "AI risk scoring in child welfare intake and investigations", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/child-welfare-referral-risk-triage. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 30 September 2026: Published after review by an automated review workflow (independent skeptic review).
  • 30 September 2026: Editorial fix pass after adversarial review: corrected the EU AI Act basis to Annex III point 5(a) (essential public services and benefits) with Article 6(3) explained as an exemption the profiling clause closes rather than a separate high risk route, and rewrote FAQ 4 to match; removed the unsupported claim that Allegheny County's evaluations were a response to public scrutiny; removed the Children's Data Network vendor entry on the Los Angeles County evidence record (the source names it as a research partner, not a builder or integrator); aligned the screening consistency quote with its source wording; rephrased FAQ 3; tightened the blitsAi.howToBuild claims on model agnostic routing and PII masking; and added the AFST's "mandatory" score band screen in rate to the human in the loop description.
  • 30 September 2026: Unpublished by an automated review workflow (independent skeptic review).
  • 29 September 2026: First published

Related use cases

Government and public sector

AI drafting of social work case notes and assessments

A generative AI tool, often built on speech to text, that turns a social worker's account of a visit or assessment, whether a recorded conversation or their own dictated or typed prompt, into a first draft of the case note or statutory assessment in the format the case record needs, for the social worker to check, correct and sign before it becomes part of the record.

Deployments
2 public, best grade B
Reported handling time reduction
63%
Swindon Borough Council, organization claim
Government and public sector

AI for benefit fraud and error detection in social security

Risk models that help a social security or benefits agency decide which claims, payments and recipients to check for fraud or error, so that caseworkers verify the riskiest cases first, while every decision on entitlement stays with a person and the model is tested for fairness before and during use.

Deployments
5 public, best grade B
Reported detection improvement
2.5x
Department for Work and Pensions, organization claim
Government and public sectorLogistics and transportation

AI for customs risk targeting, container selection and valuation checks

Machine learning models that score every import or export declaration for the risk that it carries contraband, is misclassified or is undervalued, so a customs administration sends physical inspection, scanning and detailed review to the small share of containers, vehicles and parcels that need it, while the rest clear without an officer touching them, and an officer reviews and acts on every high risk alert.

Deployments
2 public, best grade B
Autonomy
Supervised agent