Here is a number that should make you feel something you rarely feel about a federal agency. The average wait to get an answer on Social Security's national 800 number has fallen from 34 minutes to 0.6 minutes — roughly thirty-six seconds. Relief, right? Somebody fixed a thing. The number is real, it is published, it is enormous, and it is exactly the kind of figure you carry around for a week as proof that the machinery of government is not beyond repair.
Be honest about what a number like that invites you to do, because it turns out to be the subject of this piece. The number checks out (it is linked below). What it does not invite is a different sort of question — the one worth asking every time any institution tells you a wait got shorter:
Shorter for whom — measured across which people?
An average wait is not a description of a system. It is a description of the people who got through it. And in the one corner of Social Security where that distinction does real damage to real lives, the agency's own numbers now point in two directions at once, and both of them are true.
What got faster — and all of it is real
Let me put the good news down first, without qualifiers.
On August 14, 2026, the Social Security Administration marked its ninety-first birthday with a list of things it had repaired. In its own telling, it cut the initial disability claims backlog by over 30 percent, "from a high of nearly 1.3 million in FY 2024 to 884 thousand in July 2026." It fixed the phones — that 98 percent reduction from the top of this piece. And it credits itself with decreasing disability hearing wait times over 90 days, reaching historic lows.
Hold on to that last phrase. We are coming back to it.
The Commissioner said much the same thing under oath. In written testimony to a joint Ways and Means subcommittee hearing on June 10, 2026, Commissioner Bisignano reported that average processing times for initial disability claims had fallen from 233 days in May 2024 to 184 days in May 2026 — roughly seven weeks of a stranger's life given back, over two years. Representative Judy Chu called some of the agency's statistics "extremely misleading" (the one she pointed at was the phone accounting, where a caller who asks for a callback instead of holding is logged as a zero-minute wait time — not this processing-time figure). Still a reason to check; the direction survives a second source. Measured July to July instead, initial decisions came in at 186 days in July 2026, down from 220 days in July 2025. Five weeks faster than a year earlier.
Five weeks. If you have ever waited on a benefits decision while a landlord waited on you, you know exactly what five weeks is worth. I am not going to be cute about it.
One honest asterisk before we move on, because it is the agency's own arithmetic. The initial backlog did not fall steadily through 2026 — it rose, from 853,000 cases in April to 884,000 in July, and the August release does not address the three-month increase. The same analysis notes that the FY 2024 "peak" SSA measures itself against shifted by roughly 30,000 claims between two of its own publications, unexplained, and that nobody outside the building has the monthly detail that would show which factor drove what.
So far so good, more or less. Here is where it gets interesting.
The number that went the other way
At the end of the disability process sits a hearing in front of an administrative law judge. It is the stage where a human being describes to another human being what their body will and will not do anymore. It is also the stage that matters most, and I will show you why.
In fiscal 2024, the line for that hearing was the shortest it had been in a generation: fewer than 275,000 pending hearings, the lowest level in nearly 30 years (according to a law-firm study working from SSA's own releases). That same study puts the July 2025 figure at approximately 276,000, drawn from an SSA press release.
By July 2026, pending hearings had reached about 362,000.
That is more than eighty thousand additional Americans standing in a line that, two years earlier, was at a thirty-year low. And it is not one outlet's arithmetic. A legal publisher counts about 360,000 pending cases in June 2026, up from about 280,000 in October 2025. A disability-benefits explainer working from SSA's fiscal data puts the national backlog at approximately 330,000 pending cases in January 2026, up from a low of around 270,000 in January 2025. Different months, different sources, same slope.
Which brings us to the thing that trips almost everybody: SSA's claim about historic lows is not a lie. The average hearing wait really did improve — 275 days in July 2026, down from 285 days in July 2025, or about 8 months as of June 2026.
How can the average wait fall while the line gets a third longer? Because the two numbers measure different things. An average processing time is calculated over the cases that reached a decision — the people who got out. The pending count is everyone still inside. A system can move its finishers a little faster than last year while admitting far more people than it releases, and every published figure will read as progress right up until you count heads. SSA's own budget planning, per the same stage-by-stage analysis, has flagged that pending hearings could keep climbing before improving.
Which I find clarifying rather than scandalous (no one is hiding anything). The dashboard is answering a question about throughput while we read it as a question about the queue.
Where the machine sits in all this
So where is the machine? Probably not where you are picturing it.
This is not a story about a robot judge. It is not even a story about one program. Back on August 6, 2024 — well before any of these figures existed — the chairman and ranking member of the Senate Finance Committee wrote to SSA noting that the agency already employs over a dozen AI systems for tasks including reviewing medical evidence in disability claims, expediting claims likely to be awarded, and extracting handwritten data into machine-readable formats. Ron Wyden and Michael D. Crapo do not agree about much. They agreed about that.
The tool everyone names is IMAGEN. A federal technology fund bought it: the Intelligent Medical-Language Analysis Generation tool, a total investment of $1,980,560 starting in October 2024, on a page that commits to very little — SSA "intends to" improve determinations, "plans to" tackle a growing backlog, and the system is expected to give adjudicators real-time feedback, "potentially saving thousands of work hours annually."
What does it do, and where? The National Academy of Social Insurance, reading SSA's own inventory, says IMAGEN uses natural language processing and predictive analytics to organize key medical evidence, and that it is used by a portion of state disability examiners to review initial applications and reconsiderations and conduct continuing disability reviews.
Look at where that puts it, because the whole argument lives there. Initial applications. Reconsiderations. Continuing reviews. Those are the stages that consist of reading a file. Not the hearing.
At the hearing level a different program does a different job. The same task force report describes Insight as a distinct application using natural language processing "to review and identify weaknesses and inconsistencies in draft ALJ opinions," with algorithms separately used to manage case assignment. And the people who built Insight fenced it themselves, in a peer-reviewed handbook chapter: "Importantly, Insight is explicitly designed only as an assistive tool: It does not decide any element of a decision nor advise any specific remedy to potential quality issues." Two of the four authors worked at SSA; one is the creator and product owner of the software. When the person who built the thing writes that sentence in print, it beats a vendor brochure.
The third name in circulation is HeaRT, and it deserves a demotion. Hearing Recording and Transcriptions uses generative AI to produce the transcripts, does not depend on recording hardware, and saves around $5 million a year. It records the hearing. It shortens nobody's wait for a decision and it adjudicates nothing.
So put the map together. The software is deployed heavily at the paper-reading stages and, at the hearing, in a form its own designers built not to decide. Now add the one statistic that makes the shape matter. The same benefits analysis that tracked the hearing backlog puts approval rates by stage at roughly 36 percent at the initial application, 14 to 16 percent at reconsideration, and 50 to 58 percent at the ALJ hearing — with SSA's FY 2025 data showing approximately half of ALJ decisions fully favorable out of about 277,740 decisions.
Reconsideration is where most people are told no. The hearing is the first stage where approvals outpace denials. Which means the paper stages are, functionally, a conveyor toward the human stage — and the conveyor just got faster while the human stage stayed human.
Careful here, because this is where writers usually overreach. No SSA document I can reach concedes that faster upstream processing is what grew the hearing queue. The agency has not said it and I will not put words in its mouth. What I can say is that the pattern is exactly what you would expect if it were true, and nobody has offered a better account of the same numbers. An analyst-facing breakdown of the stage data draws that line explicitly; SSA, asked about criticism of its record, spoke of "streamlined processes for disability claims, smarter technology, and stronger federal-state partnerships" without naming a tool.
Now the strongest case that I have this backwards
I owe you the counter-arguments, and there are three good ones.
First: in principle this should work the other way. The same task force report that mapped where IMAGEN runs notes that identifying claims that will ultimately be allowed "can lower appeal rates and reduce the cost of later adjudications, especially at the costly administrative law judge (ALJ) hearing level." That is the design intent, and it is sound — better triage at the front end is supposed to shrink the hearing queue by approving earlier the people who were always going to be approved. If it were working that way, my thesis would be wrong. I cannot dispose of that objection. I can only note that the queue went up anyway, and that nobody has published the claim-level data that would settle it.
Second, and sharper: the front-end improvement may not be efficiency at all. Jack Smalligan at the Urban Institute looked at the same shrinking backlog and found two other things moving. Applications fell 7 percent in FY 2025 — about 163,000 fewer people — and the approval rate dropped from 38.7 percent to 36.0 percent, which by his estimate means roughly 61,000 people who would have been approved at the old rate were not. He is careful not to claim more than he knows: "It's unclear what's driving these developments, and the SSA has not acknowledged them," with "no evidence of any policy or program changes that would reduce approvals nationally." He says could be driving, and I will not upgrade him to is.
Yet his account does not rescue the system — it changes which door the crowd came through. Denied claimants are precisely the people who appeal. That last step is my inference, not his; Urban draws no line from the denial rate to the hearing backlog. But if the front end is saying no more often and faster, the stage that says yes is where those people end up.
Third: maybe the answer is more automation, not less. Mark Warshawsky, a former SSA Deputy Commissioner now at the American Enterprise Institute, traces the service crisis to manual processes, antiquated technologies and "programming languages that were once cutting edge in the 1960s but now are costly to maintain." His prescription is modernized IT and implementing new disability rules "through an automated tool" to speed administration and improve consistency. That is a serious position held by someone who has run the building.
And the independent watchdog is not on my side either. The Inspector General found the agency "has made measurable progress in improving telephone service and deploying technology to speed disability claims processing" — in the same wire story that lays out the critics' view, that the gains come from temporary staffing shifts and workforce reductions, "shifting bottlenecks around rather than solving staffing problems."
That phrase is the one I keep circling. Because something else changed at the same time as the software.
In February 2025, the agency set a target of cutting its workforce from about 57,000 employees to 50,000, a 12 percent cut, expected to come largely through retirements, resignations and voluntary separation payments. By October 2025, reporting put reductions at "more than 12 percent," bringing staffing to approximately 51,400 from 57,000 — in the same piece where SSA's September 2025 AI strategy promises to "continually reinforce humans at the center."
Fewer people, more software, faster paper stages, a longer line at the one stage that requires a person. Those four facts are not in tension. They are the same fact, told four ways.
Australia ran the national-scale version of this — and it was not AI at all
Before I go further, one correction that matters, because the internet gets this wrong constantly: Robodebt was not artificial intelligence. No model, no machine learning, nothing that learned anything. It was data-matching plus arithmetic — and that is precisely why it is the right comparison. You do not need a neural network to automate a benefits decision at national scale. You need a database and a rule.
Australia's scheme ran from July 2015 to November 2019, and the Federal Court's own summary of Prygodicz v Commonwealth of Australia (No 2) describes the mechanism in a clause: the Commonwealth took tax-office income data and evenly apportioned it over fortnightly increments in the review period, "in a process the parties called 'income averaging'." Earn your year's money in three intense months, and the computer decided you had earned it evenly across twelve, concluded you were overpaid, and sent you a bill. (If you have ever done seasonal work, you can guess how that went.)
The Commonwealth eventually admitted it did not have a proper legal basis to raise, demand or recover those debts. The court found it had unlawfully asserted debts "totalling at least $1.763 billion against approximately 433,000 Australians" and recovered approximately $751 million from about 381,000 of them. Justice Murphy called it "a shameful chapter" and "a massive failure of public administration" — then, cutting against the conspiracy reading people reach for, wrote that "given a choice between a stuff-up (even a massive one) and a conspiracy, one should usually choose a stuff up."
So why does an Australian scandal belong in a piece about Social Security?
The errors surfaced in exactly one place: the review stage. People appealed to the Administrative Appeals Tribunal, and the tribunal started finding for them. Terry Carney, an emeritus professor who spent almost forty years at the AAT, ruled Robodebt was illegal five times between April and September 2017, saying its income-averaging method lacked "sufficient strength of evidence" and "simple mathematics." Five rulings, by a single member, correctly identifying a scheme that would keep running for two more years.
What happened next is the whole lesson. The department that ran the scheme chose not to appeal his findings, which meant they were kept secret — unappealed first-tier decisions stay unpublished. Carney himself was not reappointed to the tribunal after ruling against the department. And in his later account at the University of Sydney Law School, he describes the department never appealing to the second tier any of the 220 rulings invalidating Robodebt at the first level of the tribunal. He also supplies the arithmetic that should have ended it on day one: the scheme presumed stable fortnightly casual earnings when departmental data knew this to be true for only 7 percent of clients.
Carney's counterfactual is his own, and he offers it carefully: "If my decisions had been acted on by the department, we might have had only a few hundred people who were subjected to the unlawful decision-making of Robodebt."
A few hundred, instead of the more than 500,000 victims the responsible minister referred to when the Royal Commission's report was handed down.
That report arrived on July 7, 2023. Commissioner Catherine Holmes AC SC wrote the sentence everyone remembers — Robodebt was "a crude and cruel mechanism, neither fair nor legal, and it made many people feel like criminals". But the finding this article turns on is quieter. The Commissioner counted the tribunal among the "institutional checks and balances" whose ineffectiveness "in presenting any hindrance to the Scheme's continuance" she called "equally disheartening," as Michelle Grattan reported it. The legal press put the mechanism plainly: income averaging was "inconsistent with social security legislation," and with no way to systematically review those decisions and no departmental follow-up, "the tribunal's findings were effectively ignored."
The review stage was working. It was simply invisible, slow, and easy to route around.
Fifty-seven recommendations came out of that inquiry, and the one Australia's information regulator repeats word for word should be taped to a wall at SSA headquarters: where automated decision-making is implemented, "there should be a clear path for those affected by decisions to seek review," departmental websites should explain in plain language how the process works, and business rules and algorithms should be available for independent expert scrutiny.
Australia's Attorney-General's Department then wrote down why, in a consultation paper that reads like a memo to Washington: the Commission considered that "the availability of review pathways are vital safeguards in the use of ADM," and effective merits review "is an essential part of the legal framework that protects the rights and interests of individuals."
And then — this is what separates Australia's response from a press release — they changed the machinery. The Administrative Appeals Tribunal was replaced outright: a new Administrative Review Tribunal commenced operations on October 14, 2024, all its members must identify and report system issues in administrative decision making, and significant decisions of the Tribunal will be published. Twenty-eight recommendations had been fully implemented by that point. Separately, a transparency duty for automated decisions takes effect on December 10, 2026, requiring disclosure of decisions made solely by automated processes and those "where automation plays a substantial role in supporting human decision-making."
The bill kept arriving anyway. In June 2026, three years after the inquiry, the Federal Court approved a further class-action settlement of $548.5 million for about 125,000 claimants, because the earlier 2021 settlement "did not consider information later uncovered in the Royal Commission into the Robodebt Scheme."
Let me be scrupulous: the parallel is structural, not moral. Nothing in the American record resembles what happened in Canberra. No department is burying adverse rulings; the Insight designers published their own limits; the Senate letter was bipartisan and public; SSA's tools assist rather than assert. I am not calling anybody's benefits system crude or cruel.
The parallel is narrower and, I think, more useful than an accusation. Australia learned at enormous cost that when you automate the front end of a benefits system, the review stage is the only place the errors become visible — and that a review stage which is slow, unpublished or under-resourced does the same work as a cover-up without anybody intending one. Australia's fix was not to ban the automation. It was to make the review path faster, louder and mandatory.
America's review path is the disability hearing. It currently holds about 362,000 people.
What the dashboard will say in 2032
Let me get imaginative, within the bounds of what is already deployed.
Start with the easy extrapolation. IMAGEN goes from a portion of state examiners to all of them; the medical-summarization step that eats hours of an examiner's day collapses to minutes; initial decisions fall from 186 days to 40. Reconsideration, mostly a second read of the same file, drops to two weeks. By 2032 a claim that took six months in 2026 gets its first two answers before the end of summer. This will be announced, correctly, as one of the great administrative achievements of the decade.
Now count the hearings. Administrative law judges are people. They read, they listen, they ask a woman with fibromyalgia to describe her worst day, and they cannot do it at machine speed — not from inefficiency but because that is what a hearing is. Suppose the corps grows modestly and the front end doubles its output. The queue does not grow linearly; it grows like water behind a narrowing channel.
So where does the pressure land? On the humans — because someone will notice they are the bottleneck. The proposal will not be a robot judge; nobody will propose that, and it would not survive a week. It will be modest and technically reasonable. Insight already reviews draft decisions for weaknesses and inconsistencies, and the step from reviewing a draft to producing one is small, well within current capability, and requires no announcement — because from the outside the output looks identical: a decision signed by a judge.
Then the calendar gets optimized. Cases are already grouped algorithmically for assignment; grouping them by predicted outcome is the same math with a different label. Likely denials get the twenty-minute slots. Likely approvals get waved through. Everyone is still getting a hearing. Nobody has lied about anything.
And the dashboard stays green, because every stage is measured over its own finishers. Initial decisions: excellent. Reconsiderations: excellent. Hearing processing time: improved for the eighth consecutive year. Behind that screen sits a number that gets published but never headlined — how many people are standing in the line — and a second nobody collects at all: how many stopped waiting, withdrew, gave up, or died.
No villain required — it is just what happens when you make eight steps of a nine-step process cheaper and leave the ninth one alone.
What the people who study this are saying
The most useful document I read is a phase-one report from a task force convened by the National Academy of Social Insurance, and it is useful precisely because of who sat at the table — technology companies, an insurer, a bipartisan policy shop, federal-employee unions. Not a coalition that agrees on much.
Their verdict on productivity is a model of restraint. They are "intrigued by the productivity potential" but conclude "it is premature to assume AI tools will increase worker productivity," adding that projections show the backlog will grow substantially unless agency funding and staff are increased — and noting that prominent labor economists disagree outright, with David Autor optimistic and Daron Acemoglu estimating gains will be very modest. They name the causes software cannot reach — understaffing, high staff turnover, unnecessary program complexity — and warn against "looking to AI to fix these systemic problems." Then they draw the line this whole piece has been circling: AI "will at best augment human decision making when rights-impacting benefit decisions are made," because "benefit decisions cannot be fully automated without jeopardizing the rights of a claimant."
The academics who studied these tools most closely are not hostile to them. Stanford's regulation lab frames the agency as a pioneer: despite widespread skepticism of AI in adjudication, SSA "pioneered path breaking AI tools that became embedded in multiple levels of its adjudicatory process." That is deserved admiration — and also a description of scope, which is the thing to watch.
From the Senate, the pre-commitment is already on paper, written in 2024 before any of these numbers existed. Wyden and Crapo told the agency that "AI is not a panacea for all challenges facing SSA," that SSA "must have strong governance frameworks in place that, among other important aspects, clarify the role of human discretion," and that without proper structure its use of AI "could reduce the effectiveness of its benefit administration processes, exacerbate improper payments, and jeopardize beneficiaries' financial security."
From the other direction, the Urban Institute warned in February 2025 that further staffing cuts would "exacerbate backlogs and lengthen wait times," with Smalligan and Adriana Vance predicting degraded service and "a likely increase in poverty among older adults and people with disabilities." That was a forecast. The hearing queue is now the place to check it.
And then there are the people who file these claims for a living. A qualitative study by the Disability Rights Education and Defense Fund and the American Association of People with Disabilities interviewed 52 specialists at 32 organizations and found that increased delays "affected every stage of the disability benefits process," that the new AI phone system "frequently does not interpret the caller's statements correctly," and — the line that stays with you — one advocate's read on the headline improvement: "I don't actually think it's because they're being more efficient."
SSA called that report biased; its rebuttal is the "smarter technology" line quoted earlier. Meanwhile a legal-services attorney told AARP that clients with cognitive issues "really need a person," and AARP's own testing found the phone system failed to distinguish between retirement benefits and Supplemental Security Income.
Not one of these voices argues that SSA should stop using software. Every one of them is arguing about where the human has to stay.
What does this mean for you?
If you are filing or appealing right now, understand the stage architecture. Roughly 36 percent are approved at the initial application and only 14 to 16 percent at reconsideration — but 50 to 58 percent at the hearing. A denial at reconsideration is close to the statistical norm, not a verdict on your case. Appeal within the deadline rather than starting a fresh application; a new claim puts you back at the 36 percent stage instead of the 50-plus one.
Expect your wait to be local, not national. Across 160-plus hearing offices, a wait could be as short as six months or longer than a year. The national average tells you almost nothing about yours; ask your representative or the office directly.
Treat your file as the thing the software reads. Applicants routinely submit 1,000 pages or more of medical evidence, and the tool's job is to organize and surface the key evidence inside it. Software can only find what somebody put in the folder — so chase the missing records, the specialist's notes, the functional assessment, before the file moves rather than after.
If the phone system does not understand you, that is documented, not a personal failing. It has been measured by researchers and by AARP. Ask for a person, in those words, and keep asking.
If you advocate or represent claimants, the DREDF and AAPD report is the document to put in front of a congressional staffer — it is the only one that reads this from the claimant's side of the counter.
And when any agency — any agency — tells you a wait got shorter, ask for the queue. Averages describe the people who finished. Ask how many are still in line, and whether that number is going up. It is the most useful question in this entire piece, and it costs you nothing.
The lesson, as I see it
If that phone number made you feel better at the top of this piece, hold on to the feeling. It turns out to be the point.
We judge automated government almost entirely by speed, because speed is easy to measure and easy to publish. Every number SSA put in its anniversary release is defensible. The backlog really did fall by more than 30 percent from the FY 2024 peak. Initial decisions really are five weeks faster than a year ago. Hearing processing times really did improve. Set out to falsify any of it and you would fail.
And all of it together still misses what happened: a system got much better at the steps a machine can read and no better at the step a person has to sit through — and the second kind of step is where the rights are. That is not an argument against the technology. IMAGEN organizing a thousand pages of medical records is an unambiguously good use of a computer, and I would not take it from a single examiner. It is an argument about arithmetic. Speed up eight stages of nine and you have not solved the queue; you have relocated it.
Australia paid half a billion dollars, an inquiry, and five hundred thousand people's dignity to learn that the review stage is where automated systems tell you the truth about themselves. Their answer, now sitting in Australian law, was to make that stage faster and more visible, not smaller. Ours is currently 362,000 people long, a third longer than it was a year ago, and that number is not in the anniversary release.
My vote? Publish the queue next to the average — every quarter, by hearing office, as a matter of course. Then fund the stage a machine cannot take. Because the faster the front of this system gets, the more it matters that somebody is still sitting at the back of it, listening.
If you know somebody standing in one of these lines, the useful thing to hand them is not sympathy — it is the question. Pass this along. The HAIA Foundation keeps asking what the average is an average of, and posts the answers here.






