This piece had a different headline until the last round of fact-checking, and the old headline was false.
It said that the child-welfare algorithm used in Allegheny County, Pennsylvania — a tool that has drawn years of complaints that it discriminates against poor, Black and disabled families — had cut the racial gap in who gets investigated by 83 percent. I had the study, the figure in three places that agreed, and the pleasing shape of a piece that tells you the received wisdom is upside down.
Then I did the boring thing. I opened the journal article itself and searched it for the number.
"83%" appears in it zero times. So does "10.6." So does "73%." Every figure I was about to set in a large font came out of a working paper posted in June 2023 — and the peer-reviewed version of that same study, accepted in May 2026 and published online in July, reports smaller ones.
That is the confession, and not a flattering one. I was one editing pass from putting a false number in a headline about whether a machine takes children out of their homes fairly. What saved me was a dull ninety-second habit — open the primary source, search it for the exact string, read the sentence around it. I nearly skipped it, because the number was everywhere, and everywhere is where numbers go to look true.
Here is why I am telling you rather than quietly correcting it. The federal government did not do that ninety seconds either. In March 2026 it published a brief telling states what these tools can do. The sentence about disparities carries a footnote. The footnote points at the working paper.
First, the part nobody is arguing about
A child-abuse hotline is a triage problem. Calls come in — from a teacher, a neighbor, an emergency room, an ex-partner — and a screener has minutes to decide which become an investigation and which are closed at the desk. Investigate too little and a child gets hurt; investigate too much and you send a caseworker into a family's kitchen over nothing.
In August 2016, according to Allegheny County's own documentation, the county's Department of Human Services put a predictive model into that decision. The Allegheny Family Screening Tool reads the county's integrated administrative records — prior referrals, public benefits, behavioral health, juvenile justice — and returns a risk score before the call is opened or closed. Scores above a threshold trigger a "mandatory screen-in," which sounds absolute and is not: supervisors can override the mandatory flag at their discretion, and overrides are logged and reviewed. By Washington's own account it was the first of its kind in American child protection.
The problem it was pointed at was not subtle. Before the tool arrived, referrals involving Black children in the county were 11 percentage points more likely than referrals involving White children to be screened in for an investigation. That is the gap. Everything that follows is an argument about what happened to it.
And here is the finding that made me want to write this in the first place, because it cuts hard against the expectation. Three researchers studied the switch-on and found that the tool reduced disparities in screening decisions and home removal rates for referrals involving Black versus White children — and that an analysis of false positive and false negative rates suggests this came "while improving welfare among both Black and White children." Not a tradeoff. That direction survived fifteen months of peer review, and it deserves saying plainly before I get difficult about its size.
The number, with the qualifier attached
So how big is it?
Across every referral the county fielded, the answer is 3.3 percentage points — the unconditional reduction in the screening disparity. Conditional on the algorithm's own score, it is 3.1 points. The effects, the authors note, are concentrated among high-risk referrals, the ones likely to be defaulted to a screen-in under the mandatory protocol.
Now the famous number. For referrals in the highest risk bin, the paper reports the algorithm reduced the racial gap in screening rates by 6.0 percentage points — "or 65% of the pre-existing disparity in this score bin." Read that clause twice, because it is the clause that keeps falling off. Sixty-five percent of the gap in that bin. Not of the gap in the county. Not of anything a parent would recognize as the racial gap in child protection.
Removals tell a similar story at a similar scale. Two research designs both show the tool reduced disparities in home removals within three months by approximately 1.7 percentage points — roughly 48 percent of the removal gap that existed before. Forty-eight. Not seventy-three.
The paper ran in the Journal of Policy Analysis and Management, volume 45, issue 4, published in July 2026 under an open license. Hold those numbers, because I am about to show you the other set — the ones currently doing the talking.
The sentence the authors refused to write
Before that, three things in this paper that almost nobody who quoted it repeated.
One: the authors will not tell you which direction the equalizing ran. When a gap narrows there are two ways to get there, and they are morally different: fewer Black children pulled into investigations, or more White children pulled in. The paper is explicit that its evidence cannot separate them. The change "may occur either through a reduction in the investigation rate for Black children at high risk of removal, an increase in the investigation rate for White children at high risk of removal, or some combination of these two channels." They will not say which direction it went. That is honest, and it is a hole you could drive a policy through. "The algorithm reduced surveillance of Black families" is not a sentence this study supports; neither is its opposite.
Two: the gap did not close. The authors attach important caveats the coverage skipped — chiefly that "while the AFST reduced disparities, unwarranted disparities in screening rates remained after implementation," and that "the algorithm operates only at one of many decision points in the child welfare system." A hotline screener is the front door. Everything past it — investigation, safety plan, petition, judge, placement — is untouched by this study.
And they add a warning aimed squarely at the people now holding a checkbook: policymakers considering these models elsewhere should not assume that the results will generalize without further evaluation. Their words, in their conclusion.
Three: two of the three authors built the thing they evaluated. Not an accusation — a disclosure, and it is theirs. Vaithianathan and Putnam-Hornstein disclose in the paper's own footer that they "were contracted by Allegheny County to build the AFST and continue to work with the county on other projects," and disclose interests in Social Data Analytics Ltd, "which inter alia contracts for data analytics consulting."
I want to be careful here, because this is where a piece like this usually goes cheap. Declaring a conflict is what you are supposed to do, and nothing about the disclosure makes the finding wrong. But the strongest published evidence that this tool narrowed a racial gap was produced, in part, by the people paid to build it — and that belongs next to the effect size, not in a footer you skim.
Where 83 came from
Now the part that made me rewrite this piece.
The 83 percent everyone repeated is real. It is just old. It comes from an earlier draft of the same study, dated June 2023: "for referrals in the highest risk bin, the algorithm reduced the gap in screening rates across race by 8.8 percentage points, or 83% percent of the pre-existing disparity in this score bin." Go back one more version and it moves again — a 2022 draft reports 9.6 points, "or 98% of the pre-existing gap."
So: 98 in 2022, 83 in 2023, 65 in the journal. That is not a scandal — that is what revision looks like from outside, an estimate walking down toward whatever the data will actually bear. The scandal, such as it is, happens entirely downstream.
The removal figures wandered too. The June 2023 draft's own body reports the tool cut the three-month removal disparity by 3.1 percentage points, or 77 percent of the pre-existing difference. But the figure that entered circulation was 73 percent, which you get by pulling a pre-existing gap out of one of that draft's footnotes and doing the division yourself. The published paper says 48.
And the two most quotable transitions of all — 10.6 percent down to 1.8, 4.3 down to 1.2 — appear in no version of the paper as sentences at all. They are arithmetic performed on footnotes. A footnote of unreported results in the 2023 draft gives a pre-existing coefficient of 0.106 and an effect of −0.088; subtract, move the decimal, and you have 10.6 falling to 1.8. The next footnote sums two coefficients — "0.0278 + 0.0147 = 0.0425" — and the same arithmetic gives you 4.3 and 1.2.
Where did it travel? A Forbes column on July 27, 2026 — seventeen days after the peer-reviewed version went online — carried the draft's figures as the current finding, reporting a disparity reduction "for the highest risk referral types by 83%" and a removal gap cut "by 73%, from 4.3% to 1.2%." I cite that not as authority but as the moment a stale number crosses into general circulation with the journal version already sitting there, open access and searchable.
The footnote that is now moving money
Here is where I stop being able to treat this as a media-criticism story.
In March 2026 the Administration for Children and Families — the arm of the U.S. Department of Health and Human Services that oversees child welfare — published a brief called Modernizing Child Welfare Technologies and Tools, written to be read by state administrators deciding whether to build one of these systems.
Its central claim is carefully worded, so I will quote it exactly — the whole point of this article is what happens when people paraphrase. The federal government's March 2026 brief says: "Findings from randomized controlled trials and quasi-experimental studies indicate that these tools have the potential to improve child safety, create time savings for workers, reduce disparities, and produce greater visibility and accountability for decisions."
Have the potential to. That is a hedge, and a defensible one. Keep hold of it, because almost nobody quoting this brief has.
Now follow the citation. The words "reduce disparities" carry two endnote numbers. Endnote four reads, in full: "K. Rittenhouse, E. Putnam-Hornstein, and R. Vaithianathan, 'Algorithms, Humans and Racial Disparities in Child Protection Systems: Evidence from the Allegheny Family Screening Tool' (2024)" — followed by the GitHub address of the working paper PDF.
Not a DOI. Not a journal. A file on a researcher's personal site, labeled with a year that does not match the June 2023 draft currently sitting at that address. That is the structural problem in one line: a URL like that has no version of record. Nobody is notified when its contents change, and nothing about it tells a reader in 2028 which draft the government meant. A DOI does both by construction.
The same brief holds Allegheny County up as an early adopter whose development and implementation "was guided by careful consideration of common critiques," and concludes that where these systems are implemented thoughtfully and transparently, the gains in accuracy outweigh the potential risks.
Then, on May 28, 2026, the agency announced $6 million to help state, territorial, and tribal child welfare agencies pilot new technology — a program, in its own headline, for piloting predictive analytics in child welfare.
May 28, 2026 is also the date the Journal of Policy Analysis and Management accepted the peer-reviewed version of that finding — the version with 65 instead of 83, 48 instead of 77, and the caveats.
Let me be clear about what that coincidence is and is not. It is not evidence of anything. Nobody at ACF could have known: journal acceptance is not a public event, and no alert goes out to the agencies who have been citing your working paper for a year. Two calendars landed on the same square.
But it illustrates the actual failure, which is neither dishonesty nor even carelessness. We have built a research pipeline with no return path. A working paper can travel from a personal web page into a federal brief into a funding announcement into fifty state grant applications — and when peer review revises the number downward, nothing in that chain goes back and says so. The correction sits there, open access, correct, and unread.
The strongest case against everything I have just written
Let me make the other side's argument as well as I can, because it is better than you might expect.
The direction survived. Peer review did not overturn this study; it tightened it. A 3.3-point reduction in an 11-point screening gap is not nothing — roughly a third of the disparity at the front door, achieved by software, and apparently while improving welfare among both groups.
Sixty-five is not the opposite of eighty-three. It is smaller. A brief telling states these tools "have the potential to" reduce disparities, citing a draft, would have been telling them something the published version also supports. The claim survived the revision; the magnitude shrank.
Citing working papers is normal, and often necessary. This study was received by the journal in February 2025 and accepted in May 2026 — fifteen months in review, on top of years of drafts. A government that looks at nothing until it clears peer review is perpetually citing the last decade.
And the county's own numbers point the same way. The county's own follow-up evaluation reports that after implementation, lower-risk allegations were less likely to be investigated while higher-risk ones were substantially more likely to be — described independently of the disparity paper.
All of that is true, and I hold it. Here is what it does not answer.
The failure is not that the government cited a working paper. It is that it cited a specific version of one, permanently, at an address that cannot tell you which version, in support of a sentence other people then repeat with the hedge stripped off. And the number that escaped was not the paper's headline result at all — it was the bin-specific figure, separated from the bin. Those are claims about different populations, and only one of them is in a journal.
The question that did not shrink in peer review
Meanwhile a separate question about the same tool got sharper rather than smaller.
In June 2026, three researchers published a study in the Journal of Public Child Welfare asking what the disparity paper never asked: what does this model do to disabled people? The lead author, Ian Moura, is at the Lurie Institute for Disability Policy at Brandeis, with Robyn M. Powell of Brandeis and Stetson University College of Law and H. Stephen Kaye of UCSF. The paper is titled Evaluating disability bias in the Allegheny Family Screening Tool.
Their finding, in their own words: parents and children with disabilities were significantly more likely to be penalized by all elements of the AFST scoring we examined — "socioeconomic indicators, healthcare utilization patterns, involvement with public benefits programs, and contact with the criminal legal system."
Understand why, because it is the most important mechanical idea in this story. The tool asks about no explicit disability measures. It does not need to. Disability is legible everywhere else in an American administrative record: in how often you see a doctor, which benefits you receive, what you earn, whether you have ever been picked up by police during a mental-health crisis. Feed a model those columns and you have fed it disability whether you meant to or not. The authors conclude the tool "effectively encodes disability status through proxy variables, potentially subjecting disabled parents and their children to disproportionate scrutiny and intervention," and "may reinforce rather than remedy historical biases against parents with disabilities."
Note their hedges — potentially, may — and note the method, because I am not going to do to them what was done to the disparity paper. This is a simulation. The team took the variables and coefficients from the county and applied them to three national surveys to assess whether disabled parents and children would receive worse AFST scores than non-disabled ones. They did not read real scores.
That is a real limitation, and the most famous critique of this tool carries the same one. The ACLU's own audit wrote its caveat itself: "We say 'could have' because we could not run our analysis on the actual numbers of Black and non-Black families so instead… we looked at Allegheny County data collected before the AFST was deployed to model what the risk scores would have been." It also found households with a disabled resident could be labeled as higher risk than households without one. The technical work was done by the Human Rights Data Analysis Group; the team later published its case as an overestimation of utility and risk.
So here is the fairest sentence I can write about the evidence. The strongest evidence that the tool narrowed a racial gap is a causal study of real decisions, co-authored by the people who built it. The strongest evidence that it penalizes disability is a simulation of the scoring rules, run against national survey data rather than a single real score. Each carries a structural weakness the other does not.
For completeness: in January 2023 the Justice Department took an interest, with the Associated Press reporting that it "has been scrutinizing" the tool over concerns about discrimination against families with disabilities. Scrutiny is not a finding, and none was ever announced — even the right-of-center coverage framed it precisely, as complaints to the DOJ allege. Elsewhere the argument was settled by exit: Oregon dropped its version in June 2022.
New Zealand built this first, and has still not switched it on
To see what different institutional reflexes produce from identical science, go to the other end of the world.
The intellectual origin of all this is not Pittsburgh. Work led by Professor Rhema Vaithianathan and colleagues at the Auckland University of Technology showed that integrated administrative data could identify newborn children at elevated risk of later maltreatment — the same researcher later contracted to help build the Allegheny tool. Same idea, same person, two countries, two opposite answers.
New Zealand's Ministry of Social Development commissioned in 2012 a predictive risk model that attempted to identify children at risk of physical, sexual or emotional abuse before the age of two. What happened next looks nothing like American practice.
Ethics approval came first: granted in November 2012, affirmed by the National Ethics Advisory Committee in February 2013, with oversight from He Korowai Tamariki and an advisory expert group on information security. Three reports on feasibility and ethics were commissioned before any decision about advancing to a trial, and all went through international and national expert peer review. The ministry's position is that a model should support, not replace professional judgment, and that testing would happen "in a simulated intake setting, using historical case data."
Then, in July 2015, a minister killed it. A proposed observational study that would have assigned risk scores to newborns and tracked what happened to them was blocked by Social Development Minister Anne Tolley, who said infants would not be treated as "lab rats" under her watch. Call it courage or call it cowardice; nobody can say it was done quietly.
A second effort produced the most admirable document in this entire article. The Enhancing Intake Decision-Making Project, commissioned by the Minister for Social Development in 2014, was trialed against historical intake data — and in October 2017 the agency then known as the Ministry for Vulnerable Children, Oranga Tamariki, released the documents itself under the Official Information Act. Inside them: the model referred a higher number of Māori children and young people than the status quo did, and the reason was, in the report's own word, currently unknown.
Read that again. An agency found its model disproportionately flagged Indigenous children, could not explain why, wrote that down, and released it — before deployment, in its own words.
Eleven years on, where has all that caution landed? As of March 22, 2026, when researcher Dylan A Mordaunt took stock of it, predictive modeling tools are still not used by those workers on the front line. Testing has been "carefully limited to historical, anonymised data — and carried out alongside extensive ethical, privacy and Māori-led reviews."
And I will not pretend that outcome is free. Oranga Tamariki logged more than 55,000 reports of concern in the second half of 2024 alone, a sharp increase on the year before. Those calls are being triaged right now by exhausted humans with no decision support at all, and if the Allegheny result is even directionally right, some children on the wrong side of that triage would have been better served by the model New Zealand declined to switch on. Caution has a cost.
But notice what the delay bought, because it is precisely what the United States is missing. Ethics review before code. Peer review of the feasibility case, not just the results. A published, unexplained, inconvenient finding about Indigenous children, disclosed by the agency it embarrassed. That country still has not decided whether to use this technology — and it documents what the technology does better than the country that has been running it since 2016.
Now run it forward
Let me show you where this goes.
The $6 million lands. It is not a large sum — enough for pilots in perhaps a dozen jurisdictions, which is exactly how a practice becomes a norm. Each agency writes a grant application, and each needs an evidence section. Where does that come from? The federal brief. So endnote four propagates.
Then the vendors arrive, because a federal funding line is a market signal. Slide four of the deck has the number on it. Not 3.3 — nobody sells software with 3.3 on slide four. Eighty-three, or 65 with the bin clause dropped for space, or "up to 83%," which is technically hedged and functionally a lie. A careful estimate becomes a marketing claim by losing exactly one subordinate clause.
Then the tools get built somewhere that is not Allegheny County — the part the paper's authors saw coming and said out loud. Allegheny had an unusually rich integrated data warehouse and a screening workflow with a documented override. A state without the warehouse builds from whatever it does have: Medicaid claims, benefits enrollment, criminal justice contact — three of the four categories the disability researchers identified as proxies. Thinner data does not get you a weaker Allegheny tool. It gets you a tool made almost entirely of disability and poverty signals, with the racial-disparity result attached by citation rather than by evidence.
Push out to 2030. A parent's attorney asks in a hearing why the family was investigated and receives a risk score. She asks what it was based on and gets a list of variables. She asks for the study that justified deploying it and is handed a URL — a personal web page, cited in a federal brief, updated twice since. No archived copy. No version number.
None of that requires bad faith from anyone. It requires only that every participant keeps behaving exactly as they are behaving today.
What the people who study this actually say
The useful thing about this fight is that the disagreement is not partisan.
On the civil-liberties side, the ACLU's audit did not argue from principle — it modeled the scores and said so, conceding that this is "common practice including by the county and its tool developers." The Human Rights Data Analysis Group, which does this kind of forensic statistics for human-rights work, was brought in to look at the mechanics. And the audit team's published critique argues the tool overstates both its utility and the risk it measures.
From the other end of the spectrum, the Foundation for Research on Equal Opportunity — a free-market think tank, nobody's idea of a civil-liberties shop — has published a free-market case for the same expansion the federal brief is funding. Reason's 2023 write-up, for its part, framed the disability complaints precisely as complaints.
And in the middle sit the study's own authors, who wrote the sentence that ought to be on the cover of every grant application: policymakers elsewhere should not assume these results will generalize without further evaluation.
Notice this. A free-market foundation and the ACLU disagree profoundly about whether Allegheny County should be running this tool. Neither thinks a federal brief should cite a superseded draft at an unversioned URL. That is not a left or right question. It is a competence question — and it has a fix.
So what does this mean for you?
Most readers will never see a hotline screening decision. Several of these are still yours.
Learn the ninety-second check, because it beat me and it will beat most people who quote a study at you. Find the primary source. Open it. Search it for the exact number you were told. If the string is not there, the number came from somewhere else — a press release, a draft, or someone's arithmetic.
Check the qualifier before you repeat the number. "65% of the pre-existing disparity in this score bin" is a different claim from "cut the racial gap by 65 percent." When a statistic sounds spectacular, the clause that makes it modest is usually one line above or below it. Read that line.
Check whether the citation is a DOI or a URL. A DOI resolves to a fixed version of record; a link to somebody's site resolves to whatever is there today. A load-bearing claim cited the second way has no verifiable vintage.
If your family is involved with a child-welfare agency, ask which decision the score touched. These models generally sit at the front door, and the published evidence covers only that door. Ask whether a score was used, what it triggered, and — under a mandatory-screen-in protocol — whether a supervisor overrode it. Overrides are logged and reviewed, and a logged decision is a reviewable one.
If you or someone in your family is disabled, know that the model does not need to ask. No explicit disability variable is required for a system to sort on disability, because benefits, health-care use and crisis contacts do the job. When you consent to data sharing across agencies, that is the mechanism you are consenting to. Ask what the data will be scored for, not just who will see it.
If you work for a state agency about to take pilot money, ask for three things in writing. The published journal version, not the working paper. A local validation on your own data first, because the authors themselves say not to assume the result travels. And a disability impact analysis of your specific variable set — the method is now published, and a vendor who cannot run it should not have your contract.
The lesson, as I see it
I came to this expecting a contrarian piece: everyone called the algorithm racist, the best evidence says it narrowed the gap. Some of that is still true, and the direction of the finding deserves more attention than it has had.
But the number that belonged in the headline was never 83, and never 65 either. It is 3.3 percentage points, out of an 11-point gap, at one decision point among many, with unwarranted disparities still standing afterward, in one unusually well-instrumented county, with the mechanism unresolved and two of three authors paid to build the thing. A real result — modest, contingent, carefully fenced. And not one word of it fits on slide four.
So the number lost its fences and went traveling, and by the time it reached a federal endnote it had become the reason to fund the next dozen of these. Meanwhile the finding that actually held its size — that a model can encode disability through proxies without ever asking about disability — is still waiting for anybody in Washington to footnote it.
Peer review is the closest thing public knowledge has to version control. It did its job here, walking a number from 98 to 83 to 65 and attaching the caveats that make it usable. What we never built is the part that tells everyone downstream when the version changed. Until we do, the fastest-moving number in any argument will be the earliest one — the biggest, roundest, least qualified draft.
Cite the version. Keep the clause. And when a number looks too good to be true in a field where the outcome is whether a caseworker knocks on a family's door, spend the ninety seconds. I nearly didn't.
Ninety seconds with a search box beat me to my own headline. The HAIA Foundation reads the footnotes — and the weekly version of that lands over here.






