There is a sentence I have been quietly reassured by for years, and until this spring I had never turned it over to see what was underneath. It appears in vendor brochures, in procurement answers, in the trust-and-safety corner of every company that sells software which decides something about a person: we test our systems for bias.
I read that line the way you read a health-inspection grade taped inside a restaurant window — as a claim somebody whose job it was to check had already checked. An audit, in my head, came with a reader attached. Somebody receives it. Somebody can be embarrassed by it. That is the entire reason the word carries any weight.
What I never did was ask who that somebody was. And when I finally asked — because a federal magistrate judge in San Francisco spent thirteen pages answering exactly that question — the answer, on the record in front of her, was the company's own lawyers, and nobody else. That is not a scandal; as far as I can tell it is a careful application of a very old rule in a narrow, honest opinion. It is also the most consequential thing to happen to algorithmic accountability this year, and it happened in a discovery order almost nobody will ever read.
First, the part nobody disputes: what the judge was actually asked
The case is Mobley v. Workday, and the docket has been open since February 21, 2023, assigned to Judge Rita F. Lin and referred to Magistrate Judge Laurel D. Beeler in the Northern District of California. The plaintiffs, in the court's own summary, are "suing Workday, Inc., for utilizing an AI screening system that is more likely to deny applicants who are African American, suffer from disabilities, or are over forty years old." Allegation, not finding — hold that thought, because it matters all the way through.
Derek Mobley says the platform's algorithms caused him to be rejected from more than 100 jobs over seven years on account of his age, race and disabilities; four other plaintiffs later joined on age. CNN reported the platform is used by over 11,000 organizations worldwide, with millions of open jobs listed on it each month — which is why this reaches a long way past one man's inbox. Workday's answer, given in May 2025, was that the ruling then in the news was "a preliminary, procedural ruling … that relies on allegations, not evidence," and that "We continue to believe this case is without merit." Hold both things at once: the claim is unproven, and it is live.
Two rulings got it this far. Mobley does not allege Workday was his employer; he alleges it may nonetheless be held liable as an "agent" of the employers using its software — and in July 2024 the court let that theory proceed, denying Workday's second motion to dismiss. A theory that survived a motion to dismiss, not a finding that the vendor is liable. Then, on May 16, 2025, Judge Lin conditionally certified the age claims as a nationwide collective, writing that applicants "are alike in the central way that matters: they were allegedly required to compete on unequal footing due to Workday's discriminatory AI recommendations." When the defense warned the collective could run to hundreds of millions, her answer was that "Allegedly widespread discrimination is not a basis for denying notice."
So the case survived, and everything turned on discovery. Reporters saw that coming: a year before the ruling I am about to describe, Bloomberg Law was already calling the fight over documents a pivotal moment for AI bias lawsuits, with the plaintiffs seeking "results of internal audits on its hiring tools, and technical details about how its algorithms operate," and Seyfarth Shaw's Andrew L. Scroggins, for the defense, calling the request "such a gigantic bite." Why does discovery decide everything here? Because, as that same reporting put it, these claims are "difficult to prove given job applicants' lack of insight into how AI evaluates them." You were not told the software rejected you. The only people who know how the system performs across demographic groups are the people who tested it.
Which brings us to the testing.
The sentence the ruling turns on
On an order signed May 28, 2026 and docketed the next day, Magistrate Judge Beeler denied the plaintiffs' motion to compel production of Workday's bias-testing data. Her holding: "Attorney-client privilege applies to the bias-testing data because Workday's attorneys curated the underlying data and used the results in providing legal advice." Duane Morris, whose class-action team posted the opinion, gives the formal citation as 2026 WL 1510537, ECF No. 340.
Read the sentence the whole thing turns on, because every clause in it is doing work:
"Here, Workday has represented that its attorneys curated the data it used in the bias testing, the overall purpose of the testing was to provide legal advice and not to be used in a business capacity, and it has not submitted the data to a regulatory body."
Three representations. Lawyers assembled the inputs. The output was legal advice rather than a business document. Nothing went to a regulator. Satisfy all three and the most important self-knowledge a company has about its own hiring system sits behind the oldest privilege in Anglo-American law.
Note what the ruling does not say. The plaintiffs argued Workday had waived the privilege by pointing publicly at the fact that it tests — they cited an "AI Fact Sheet," in the court's words, "which states that Workday performs bias testing." The court's answer is the most quotable line in the order: "Workday's invoking the mere existence of its bias testing outside of litigation is not enough to waive privilege."
Saying you test is not the same as showing what you found. Those can now be two entirely different documents, and only one of them has a reader outside the building.
To be fair to the court — and this matters, because the ruling is being summarized more broadly than it deserves — Judge Beeler fenced it in carefully.
She did not hold that bias testing is beside the point. She held the opposite: "Workday's relevance argument is unpersuasive. Even assuming that Workday's bias-testing data is not the best evidence of disparate impact, it has at least some probative value." Relevant, and shielded anyway. She bounded the holding to what was actually claimed: "there may be some information that is not privileged…This order addresses only the bias-testing data that Workday is claiming is privileged." Norton Rose Fulbright's write-up confirms the practical shape of that — Workday had already produced the underlying historical data used in its bias audit. The raw material was handed over; the analysis was not. And she restated the doctrine that cuts against reading her own order expansively: "Because it impedes full and free discovery of the truth, the attorney-client privilege is strictly construed."
The plaintiffs did not lose everything that day, either: "The court orders production of Workday's EEO-1 and OFCCP documents. They are relevant to Workday's knowledge of potential demographic disparities when utilizing the AI tools." They did lose a second fight for an entirely different reason — the applicant data held by Workday's customers stayed out of reach, and it is still out of reach, though not for the reason the order first gave. Judge Lin vacated that piece of it on June 24, 2026 and sent the question back, and on July 12 the magistrate denied production again on a different ground: "The plaintiffs have not met their burden of establishing possession by Workday of its customers' data."
Even the firms advertising this ruling to clients keep the hedge: Norton Rose Fulbright's headline claim is that bias-testing data "may be protected" against discovery — not is, not always. One magistrate, one record, one discovery dispute. Anyone telling you courts have now ruled AI bias audits are secret is selling something.
So far so good. Here is where things get interesting.
Now put that beside the rules telling employers to grade their own homework
The reason this discovery order outgrows its own docket is that the entire American approach to regulating automated decisions rests on a self-assessment — and the assessments are arriving right now.
Start with California, which has the most developed regime. The California Privacy Protection Agency finalized its rules on automated decision-making technology, risk assessments and cybersecurity; they took effect January 1, 2026, and the deadline for businesses using automated decisions in hiring is written into the regulation itself: "A business that uses ADMT for a significant decision prior to January 1, 2027, must be in compliance with the requirements of this Article no later than January 1, 2027."
So what does California actually receive? Count the items in the submission section and the answer is bracing: the number of risk assessments the business conducted, plus an "Attestation to the following statement: 'I attest that the business has conducted a risk assessment…Under penalty of perjury under the laws of the state of California, I hereby declare that the risk assessment information submitted is true and correct.'"
A headcount and a sworn sentence. The substance — the actual report — is available to exactly two readers, and only if they ask: "The Agency or the Attorney General may require a business to submit its risk assessment reports…A business must submit its risk assessment reports within 30 calendar days." To the public, never. The first paperwork does not even land in Sacramento until April 1, 2028, and there is a small carve-out written in for lawyers: the assessment must identify everyone who provided information for it, "with an exception for legal counsel who provided legal advice."
I searched the full 127-page approved text for "attorney-client," "work product," "waiver" and "Public Records Act." Zero hits — and I want to be precise, because negative findings get overstated. The regulation does not deny privilege and it does not grant it. It simply does not address the question.
California's employment regulator came at the same problem from a different direction. Its rules on automated decision systems took effect on October 1, 2025 and require employers to preserve automated-decision records — dataset descriptors, scoring outputs, audit findings — for four years. Note what they do and do not do: they are not a testing mandate. They make testing evidence. "In discrimination cases, courts and agencies may consider the quality, scope, recency, results, and employer response to bias testing. The absence of such evidence may weigh against employers that choose not to evaluate their ADS."
Read those two developments together and tell me what a rational, well-advised employer does. Do not test, and the absence may weigh against you. Do test, and the results are relevant evidence of disparate impact — the Workday court said so explicitly. Exactly one configuration gets you the credit without the exposure, and the entire American bar worked it out inside a week. Read what employers are being told in plain language: "An audit commissioned by HR or a business unit, without counsel's involvement, is far more likely to be deemed a routine business record than a privileged communication." That is not a loophole found by a rogue firm. That is competent lawyering, published openly, in August 2026.
Meanwhile Colorado — which everyone still cites as the state that mandated impact assessments — went the other way entirely. On May 14, 2026, two weeks before the Workday order, Governor Jared Polis signed Senate Bill 26-189, which "repeals and reenacts" the 2024 AI Act as the Automated Decision-Making Technology Act. Read Skadden's summary slowly: the new act "eliminates the discrimination-related obligations" of the old one, "such as the duty of care to mitigate algorithmic discrimination risks, requirements regarding algorithmic discrimination, and requirements to perform annual impact assessments and maintain a risk management program." The effective date slid from June 30, 2026 to January 1, 2027 along the way. What arrives in 2027 is a disclosure-and-human-review framework whose details are still being written as I type: the Colorado Attorney General's rulemaking is open, and comments are being taken through October 26, 2026.
And the federal backstop? Quieter. The law has not changed — federal disability guidance still says an employer can be responsible "when an employer uses another company's discriminatory hiring technologies." The explanations changed. After a leadership change at the Equal Employment Opportunity Commission, its technical-assistance documents on AI were removed from the agency's website, even though, as the firm that catalogued the removals noted, "Federal, state and local anti-discrimination laws still apply to AI use." Then, on April 23, 2025, an executive order titled "Restoring Equality of Opportunity and Meritocracy" told federal agencies to "deprioritize" enforcement of statutes carrying disparate-impact liability. Private plaintiffs and state law are another matter — that liability "remains codified in Title VII and certain other state and local statutes" — which is exactly why a private discovery fight in San Francisco now carries so much freight.
So that settles it — the judge broke algorithmic accountability. Right?
No. And the honest version of the counter-argument is strong enough that I want to make it properly before arguing with it.
Attorney-client privilege is not a tech-industry invention. It exists because a client who cannot speak candidly to a lawyer gets worse legal advice, and a society with worse legal advice gets worse compliance. If a general counsel cannot say "run the numbers and tell me honestly whether we have a problem" without that inquiry becoming Exhibit A eighteen months later, the rational instruction is not to ask. On this view the privilege is not what hides the bad news; it is what causes anyone to go looking. Judge Beeler applied the doctrine narrowly, ordered the EEO-1 filings produced, and did not bless secrecy as policy. We have been here before, too: a long list of states enacted environmental audit privilege and penalty-immunity laws on exactly this theory — promise a company its self-inspection will not be used against it, and you get more self-inspection. Whether that bargain paid off is genuinely contested. It is not a crank position.
There is also a serious argument that the mandates themselves are the problem. The R Street Institute ranked "'AI Fairness' & Anti-Bias Laws" first among the worst state AI policies — a report published February 12, 2026 by Adam Thierer and Logan Kolas, faulting regimes that generate "mountains of paperwork, tickytack compliance regiments, unclear reporting requirements." If the assessment is a ritual producing a PDF nobody reads, shielding it from discovery costs the public approximately nothing, and the real reform is to stop demanding the PDF.
I take that seriously. I just think it points somewhere its authors would not like: if the paperwork is worthless because nobody reads it, the fix is a reader — not less paper. Which is not hypothetical. One jurisdiction tried it.
New York City went the other way: make the audit public or do not use the tool
Here is the detail that reframes the entire case, and it has been sitting inside the order the whole time, recorded as a plain matter of fact:
"In 2023, Workday engaged an external consultant to conduct testing of its product Spotlight using the impact ratios approach described in NYC Local Law 144."
That test was not claimed as privileged — the firm write-up quoted earlier says the same — and the court rejected the argument that disclosing it waived anything else: "plaintiffs have not shown that a waiver of bias-testing data in one context (here, by disclosing the HiredScore testing data to auditors) warrants a waiver of all bias-testing data."
Two tests, one company, two completely different fates. The difference was not the mathematics and not the technology. It was who the test was for. The one performed for a public-disclosure regime stayed in the open; the one performed for counsel went behind the shield.
So what is the regime that produced the open one? Under New York City's Local Law 144, an employer may not use an automated employment decision tool unless it has been bias-audited and "information about the bias audit is publicly available." The city's consumer-protection department began enforcing the law and rule on July 5, 2023, making New York, as one research team put it, "the first jurisdiction globally to mandate bias audits for commercial algorithmic systems."
Now read the mechanism, because this is the part the other regimes do not have. The adopted rules require the employer to publish the audit summary on the employment section of its website, before use, "in a clear and conspicuous manner." Not filed with an agency. Not attested to under penalty of perjury. Posted — where an applicant can find it. And the rules are specific about the contents: the date of the most recent audit plus "the source and explanation of the data used to conduct the bias audit, the number of individuals the AEDT assessed that fall within an unknown category…the selection or scoring rates, as applicable, and the impact ratios for all categories." It stays posted at least six months past the tool's last use. The auditor must be genuinely independent — not someone "involved in using, developing, or distributing the AEDT," with no employment relationship and no direct or material indirect financial interest. There is an escape hatch, and even that is disclosed: a category under 2% of the data may be excluded, but "the summary of results must include the independent auditor's justification for the exclusion."
Notice what that does that no privilege doctrine can undo. It never asks whether the audit was commissioned by counsel or by HR. It makes publication a condition of use. The audit has a reader by construction, and the reader is the person the tool is pointed at.
So does it work? Here I have to be honest with you: not nearly as well as it should.
New York State audited New York City's enforcement of its own law. The State Comptroller's report, issued December 2, 2025 and covering July 2023 through June 2025, found "DCWP's AEDT complaint process is ineffective in ensuring that all complaints related to non-compliance with LL144 are routed to DCWP" — then put a number on the gap: "DCWP surveyed websites and bias audits of 32 companies and identified just a single issue of non-compliance…We reviewed the same companies and identified at least 17 instances of potential non-compliance under LL144."
Same 32 companies. Same public documents. One finding versus at least seventeen. The penalties were never the constraint — employers face up to $1,500 per violation per day, and DLA Piper expects "a new phase of stringent enforcement." Attention was the constraint.
The research is bleaker still. One team went looking for the posted audits the way an applicant would, sending 155 undergraduates from Cornell's Communication and Technology class out in a partnership across Cornell's Citizens and Technology Lab, the Data & Society Research Institute and Consumer Reports. Out of 391 employers, 18 posted audit reports and 13 posted transparency notices telling job-seekers their rights. The researchers are careful in a way I want to copy: because the law "grants employers substantial discretion over whether their system is in scope," a missing audit "cannot be said to indicate non-compliance." Absence is not proof. It is just absence.
A second team ran the numbers a year later and found the same shape, collecting "44 bias audit reports covering 116 bias audits" — audits for roughly 2% of the Fortune 500, "despite an industry report finding more than 98% of Fortune 500 companies use applicant tracking software." Publishing an audit does not automatically make it a good audit, either: the reports are "significantly hampered by several issues, including missing demographic data, opaque data aggregation, problematic uses of 'test data,' and reliance on metrics that do not represent how automated hiring tools are used in practice." Their carefully hedged conclusion is that these tools "could often be in violation of the four-fifths rule when considering potential impacts of missing demographic data." Could. Not do.
Anyone who tells you Local Law 144 is a success has not read the Comptroller.
Here is why I still think it is the better architecture: every criticism above was written by an outsider. The Comptroller could count violations because the documents were public. The students could measure compliance because the answer was supposed to be on a careers page. Researchers could grade 116 audits because 44 reports existed to grade. A weak audit you can criticize produces more accountability than a rigorous one nobody may see — the weak one can be improved by people outside the company; the rigorous one only by people inside it.
Now run the clock forward
Just imagine it is 2031, and the incentives above have simply been allowed to run.
Every serious employer tests its hiring systems, because the alternative — an empty file where the evaluation should be — is a liability in itself. The tests are thorough, well-designed and quarterly. Not one is commissioned by a hiring manager. They are all commissioned by counsel, structured as legal advice, with a lawyer listed as the requesting party on the first page, because that is what every competent firm advises and the advice is correct. The compliance market has adjusted around it: "privileged algorithmic assessment" is a product category with a pricing page, and the workflow software routes the engagement letter through the legal department automatically. There is a tidy public artifact — a trust page, a line in the annual report about systems being regularly evaluated for fairness — and a sealed one containing the numbers. Both are true. Only one is a document anybody can ask for.
Then the pressure point arrives, and it is the third clause of Judge Beeler's sentence: the privilege held in part because Workday "has not submitted the data to a regulatory body." California can demand a risk assessment within 30 calendar days of asking, and the regulation says nothing about privilege. So the first time an agency actually asks a large employer for the underlying report, somebody must decide whether handing it over blows up the shield in every private lawsuit for the next decade. My guess — and it is a guess — is a negotiated, heavily lawyered, minimally responsive submission, followed by a quiet campaign for an explicit statutory audit privilege modeled on the environmental ones. In several states it passes. It sounds reasonable. On its own terms it is reasonable.
Meanwhile the audits that do exist get aggregated in the way that flatters. Stanford researchers have shown how that trick works without anyone intending it: pool every job together and opposite-direction disparities cancel out, so a system that disadvantages one group in warehouse roles and another in clerical roles reports clean numbers overall.
And the applicant? She sends out a couple of hundred applications a year and is filtered out of most of them by a model she never sees. In New York City she can, in principle, find a posted summary with impact ratios on it — if her employer decided the tool was in scope. Everywhere else, her case requires her to prove a pattern using documents that exist, that are relevant, that a court agrees have probative value, and that she cannot have.
Nothing in that is illegal or even underhanded. It is what happens when you require a measurement, attach a consequence to the result, and leave the question of who gets to read it to the ordinary operation of privilege law.
What the smart people are saying
The strange thing about this ruling is how thoroughly it was anticipated.
In the Harvard Journal of Law & Technology in Spring 2021, Ifeoma Ajunwa — then at the University of North Carolina School of Law — argued that trade secret law "protects automated hiring systems from outside scrutiny and allows discrimination to go undetected." Her claim was stronger than audits are a good idea: "I argue not just that the law allows for the audits, but that the spirit of antidiscrimination law requires it." She was writing about intellectual property as the shield. Five years later the shield is privilege — older, better established, much harder to legislate around.
Ajunwa was building on work by Pauline Kim, the Daniel Noyes Kirby Professor of Law at Washington University School of Law, whose research centers on discrimination risk when AI makes consequential decisions in employment, housing and credit. As Ajunwa characterizes Kim's position, auditing should be an important strategy for examining whether automated hiring outcomes comport with equal opportunity guidelines — exactly the strategy this ruling makes optional to disclose.
There is also a name for an audit that exists mainly to be pointed at. Writing for the German Marshall Fund in November 2022, Rutgers law professor Ellen P. Goodman and Julia Tréhu called it "audit-washing": "Inadequate audits or those without clear standards provide false assurance of compliance with norms and laws." Their underlying diagnosis should worry every legislator drafting an assessment mandate today — the "algorithmic audit" itself "remains ill-defined and inexact."
From the civil-rights side, the argument has been on the federal record since January 31, 2023, when ReNika Moore, director of the ACLU's Racial Justice Program, told the EEOC that "we must have comprehensive public oversight, transparency, and accountability" so that jobseekers "do not face the same old discrimination dressed up in new clothes." She named the asymmetry precisely: many hiring technologies "are invisible to workers," which "has made it more challenging for individuals and the private bar to file complaints." You cannot complain about a decision you did not know was made.
R Street's objection belongs in the same breath, because the two critiques share a diagnosis and split only on the cure: define the assessments better and publish them, or stop requiring them. Pick a side honestly — but nobody's first choice is the status quo, which is mandatory, undefined and sealed.
And the underlying harm is not speculative. Researchers at Stanford's Institute for Human-Centered AI — Rishi Bommasani, Sarah H. Bana, Kathleen A. Creel, Dan Jurafsky and Percy Liang — found "substantial evidence of racial disparities in AI-based candidate screening," measured with the EEOC's own four-fifths rule: "26% of Black applicants and 15% of Asian applicants applied to positions where the AI system discriminated against their racial group." Roughly one in four, in a market where nobody is told which systems those were.
What does this mean for you?
Not much of this is in your hands. Some of it is, and those parts matter more than they look.
If you are job hunting in New York City, go and look. The summary belongs on the employment section of the employer's website, carrying impact ratios for all categories and the date of the most recent audit. Two minutes before you apply. Finding it tells you something; not finding it tells you something too — though, per the researchers, not proof of anything.
Ask the question anyway, everywhere else. After a rejection, or on a screening call: was an automated tool used to evaluate me, and has it been audited for disparate impact? You will usually get a non-answer — note its shape, because "our vendor handles that" and "we cannot share that, it is privileged" are very different sentences.
Keep your own file. Dates, roles, platforms, outcomes. A pattern across a hundred applications is the only evidence a person in your position can generate — and in cases like Mobley, the named plaintiff's own record is where it started.
If you live in Colorado, the comment window is open until October 26, 2026. The Attorney General is writing the rules that govern automated decisions in your state from January 1, 2027, and whether they require anything to be published is being decided by whoever shows up.
If you live in California, know what your state actually gets. A count and an attestation; the report only on request. If you think "publicly available" belongs in that regulation the way it does in New York City's, your representatives are reachable and the comment record is real.
If you buy or run this software, see the vise. Test it — in California the absence of testing evidence may weigh against you, and records must be kept four years. But recognize the choice about who commissions the test, because that choice, not the result, decides who is ever allowed to read it.
Do not over-read the ruling in either direction. One magistrate judge, one record, one discovery dispute, expressly limited to the data one company claimed. It is not a holding that bias audits are secret.
The lesson, as I see it
We have quietly built an accountability system with a hole in the middle. Every serious proposal for governing automated decisions asks the deployer to measure its own system and write down what it finds — a sound instinct, since nobody else has the data. But almost none of them answered the follow-up question, which is not will the measurement happen but who is entitled to read it. California answered "a regulator, if it asks." Colorado deleted its answer. The federal government stopped explaining itself. Into that silence walked the ordinary, correct, centuries-old operation of attorney-client privilege, which answers the question the way it answers everything: the client and the lawyer.
The word "audit" is doing two jobs now, and they are not compatible. One produces knowledge that changes a decision. The other produces a defense. You can tell them apart with the single question I failed to ask for years: who reads it?
My vote? Put a reader on it. Not necessarily the world — a regulator with the staff to look, a researcher under a data-use agreement, an applicant with a link on a careers page. New York City's version is under-enforced, under-complied-with and imperfectly designed, and it still produced 44 public reports outside experts could grade, a state audit that caught 17 problems the city missed, and 391 employers whose careers pages a stranger was entitled to check. A low bar — and, at the moment, the highest one anybody has cleared.
An audit with no reader is a private opinion with a methodology. HAIA goes looking for the ones a stranger is allowed to read, which is a shorter list than it should be — the results land here.






