The pictures at the top of these articles are not photographs. Every one was made by a model, and I have always assumed the file says so — not to you, but to any software that thinks to ask. There is a field for exactly that: when a generator writes IPTC's trainedAlgorithmicMedia value into an image, the file states, machine-readably, that no camera was involved.
This week I checked, because a California statute operative since August 2 rests on the premise that statements like that stay where you put them.
They do exist. The file in my publisher's storage bucket carries the field, correctly set, beside a plain-English credit line naming the tool. Then I fetched the same image the way your browser does when you open one of my own posts — through the resizing URL the published page renders — and ran the same command. The credit line survived. The machine-readable field was gone, with the whole block it lived in. Stored: 1,319,495 bytes. Served: 403,711. What gave was the part machines read.
So, the confession: I have written about provenance for months while assuming my own pipeline preserved it, and never once ran the command that would have told me otherwise.
The fences, because this gets quoted badly. Substack is not a covered provider under the California law, and the duty telling large platforms not to strip provenance does not attach until January 1, 2027 — nothing here violates anything. It is one publisher's pipeline, one day, images made through somebody else's API: a mechanism demonstrated, not a study. And nobody was sneaky. Resizing for the web discards metadata, because that is much of what resizing is for.
Which is the problem in miniature. The label is not being attacked; it is being optimized away by software nobody told it mattered.
The statute is real. It is also much smaller than its own headline.
California's AI Transparency Act began as Senate Bill 942 in 2024 and reached you amended. Its start date is not in dispute if you read the code — the 2025 amendment that pushed it to August 2, 2026 says "This chapter shall become operative on August 2, 2026" — and mildly in dispute if you read the press releases, since the author's own office announced that the law took effect August 1 and local coverage said only that a new law took effect over the weekend. You will also still see January 1, 2026 quoted — true until the amendment moved it, and, as one tracker puts it, still cited across much of the web. Practitioners call the new date alignment with the EU AI Act, whose transparency article became applicable the same day. Their explanation, not the bill's.
Who is bound? A covered provider: an outfit producing a generative AI system with over 1,000,000 monthly visitors or users, publicly accessible in California. Below that line the statute does not impose a lighter duty. It imposes none — which is where the headline version of this law, including the one I started with, goes wrong. There is no optional label down there, because there is nothing down there.
For everyone above the line, three things arrived on August 2.
A free detection tool. It must accept a file or URL and expose an API. But read the scope: it tells you whether content was created or altered by that provider's own system, and nothing else. Twenty compliant tools would answer twenty different questions, none of them yours — did a machine make this?
A visible label — the optional one. A covered provider "shall offer the user the option to include a manifest disclosure," a non-hidden mark. The provider must offer. The user decides.
An invisible label — the mandatory one. The latent disclosure is provenance data written into the file itself, for software rather than eyes. The part with teeth, and the part nobody photographs.
Here is where it gets interesting, and this is the most important paragraph here. The embed itself is not hedged: a covered provider "shall include" a latent disclosure. What that disclosure has to convey is required only "to the extent that it is technically feasible and reasonable," and its durability — "permanent or extraordinarily difficult to remove" — carries the same words again: to the extent it is technically feasible. And when the duty reaches large platforms in 2027, they "shall not, to the extent technically feasible, knowingly strip" compliant provenance data from uploads. Three requirements on one signal, the same qualifier bolted to each. Not a loophole — legislatures write it because engineers testify that some things are not feasible. But the regime's strength is then set by whatever comes to count as feasible, and my resized image suggests today's answer is whatever the pipeline already does.
Three facts the coverage drops. The calendar holds three separate deadlines and only the first has arrived: platforms and download sites in 2027, camera makers in 2028. Penalties run at $5,000 per violation, with each day a discrete violation — serious money — but the act does not create a private right of action, so it moves only when a public lawyer moves it. And the word "text" was removed from the substantive provisions, so the duties reach image, video and audio only. Every paragraph a chatbot writes for you sits outside this law.
The best argument against this law is that the technology cannot do the job
And it was made loudly. In April 2024 a TechNet-led coalition — with the California Chamber of Commerce, the Computer & Communications Industry Association and NetChoice — told legislators the bill "requires platforms to comply with technically infeasible and impossible standards," and that watermarking is "still incredibly unreliable and in many cases easy to break." Advocacy, not a finding, but still the sharpest technical objection on the record. Then, that August, the industry coalition agreed to formally withdraw its opposition after negotiating with the bill's author — which suggests the fight was about cost and timing more than physics. Chamber of Progress kept pushing, urging a veto of a bill requiring "technology which simply does not exist" and, in the letter's words, "may never exist." Newsom signed it anyway.
The independent evidence is not kind either. A NeurIPS paper titled, flatly, Invisible Image Watermarks Are Provably Removable Using Generative AI found that adding noise to an image and reconstructing it stripped 98% of invisible marks under one resilient scheme, RivaGAN, while holding quality above a defined threshold — a result about that scheme under those conditions, not a universal 98%, and bad enough as stated. An audit of image generators found that, as of the study, only a minority — 38% — implemented adequate watermarking; a status quo, not anyone's compliance, since the duties came later. Removal is cheap, too: a free tool stripping AI watermarks from Claude's text appeared on GitHub within 24 hours of Anthropic's watermarking feature going live — text, which California does not cover, though the economics generalize. And a hit is not proof of authorship: Anthropic's framing is that a match shows its model "may have processed" the content, not that it wrote it.
Then the ordinary-user problem. Hany Farid has spent a career on image forensics, and asked how he verifies one company's watermark, he explains that he can do it only because his firm has a relationship with that company — adding, of telling real from fake by eye, "I can barely do it reliably, and this is what I do for a living." The best-known public detector portal, Google's SynthID Detector, was announced in May 2025 inviting journalists, media professionals and researchers to join our waitlist — and it scans for that company's own mark.
The schemes also do not talk to each other. In the same session I checked 662 AI-generated images in this repository. Files from one major model carried the IPTC field and no C2PA manifest; files from another carried a full C2PA manifest — a signed, structured provenance record — and no IPTC field. Two serious generators, two incompatible ways of saying the same sentence.
So is the law theater? I don't think so. It never promised a lie detector; it asks for a signal that travels, and the weak point in a traveling signal is plumbing, not cryptography. The World Privacy Forum's review of C2PA lands there: strip metadata when files are published or distributed and "the provenance chain — and therefore the promise of tracing content provenance throughout the digital media ecosystem — is disrupted." In e-discovery, a native image may carry its disclosure through collection and hashing, then lose it in processing. No villain required. Just a resizer.
Seoul got there first, and wrote the rule without a number in it
On January 22, 2026 — just over six months before California's date — South Korea's AI Framework Act took effect, the first comprehensive national AI statute anywhere to be enforced. Generative outputs must be labeled as AI-generated, realistic synthetic audio, images and video clearly disclosed, and the law reaches conduct abroad affecting Korean users.
The interesting part is what its drafters did not do. Article 31, the labeling provision, holds no size test: where content is a deepfake — "virtual sounds, images, videos, etc. that can be mistaken for real" — the provider "must clearly indicate the fact that the work has been generated using AI."
Korea does use thresholds; read the U.S. Commercial Service's own briefing and they sit elsewhere — on frontier-scale compute, and on which foreign operators must appoint a domestic representative. Here is the irony I cannot let pass: one of those numbers is one million users — daily users, over three months — and it decides who needs a representative in Seoul, not who must label anything. Both jurisdictions reached for the same figure. Only one attached it to the duty that protects readers.
Don't oversell it, though: Korea's regime has gaps, just not size-shaped ones. A week in, businesses still did not know when a watermark was needed or who exactly must apply it, and firms that merely use generative tools — animation and webtoon studios among them — are not AI service providers and fall outside the labeling duty. Threshold-free is not gap-free. Industry objected in advance, with The Korea Times reporting that firms called the mandatory watermarking rules "particularly worrisome", and on day one the Korea JoongAng Daily noted an overseas app advertising an "AI eraser" that had already passed half a million downloads.
Set enforcement side by side. Korea covers essentially every AI business operator serving its users — the Future of Privacy Forum reads the definition to include corporations, organizations, government agencies and individuals — and caps administrative fines at KRW 30 million, roughly twenty thousand dollars, with a grace period easing inspections. California covers almost nobody and charges $5,000 for every day of non-compliance. Broad and gentle against narrow and sharp. I do not know which is the better bet, and anyone who says it is obvious is selling something.
China points the same way: in China Law Translate's rendering of measures effective September 1, 2025, the scope article names no user count either, and distribution platforms were bound from the start, where California defers that by sixteen months.
Just imagine the next three deadlines
None of this needs a breakthrough. Only the calendar already written into California law.
Imagine January 2027, when the platform duty lands. Every large service must stop knowingly stripping compliant provenance — to the extent technically feasible. But every large service is, functionally, a resizing pipeline, and carrying metadata through one is an engineering project, not a physical impossibility. So the question becomes a budget conversation at a few dozen companies, settled quietly, with no private right of action to drag it into daylight. The same month, sites making model weights available for download acquire duties too — the open-weight pipeline, five months behind the providers.
Imagine January 2028, when it reaches cameras. Devices sold in California start offering owners a latent disclosure, and the game inverts: once signed authenticity is normal, an unsigned photograph starts to look suspicious — and the people holding unsigned photographs are those with older phones, stripped uploads, and reasons to stay anonymous.
Imagine 2029, and the two-tier internet. Above the line: labeled, embedded, checkable. Below it nothing is required, and that population is not small. Mapping 31 nonconsensual deepfake sites, the Institute for Strategic Dialogue found May 2025 traffic ranging from 375 visits to nearly four million — a long tail, many of them orders of magnitude below a million, though nobody has counted how many sit under California's line. The content most engineered to fool you is not coming from the companies most likely to be covered.
What the people who study this actually argue about
The disagreement does not sort by ideology, which is the surest sign it is real.
From the market side, the Center for Data Innovation has argued since 2024 that watermarking may not be feasible or effective for some types of media and, on images, is not a foolproof solution against deepfakes. Free-speech lawyers object differently: the Foundation for Individual Rights and Expression argues that government-compelled speech — "whether that speech is an opinion, or fact, or even just metadata" — is generally anathema to the First Amendment. Which is why labeling is the cautious option: surveying election-deepfake laws, the R Street Institute counted 17 of the first 20 states choosing a labeling requirement over prohibition, after a federal judge blocked California's prohibition approach in 2024 on First Amendment grounds.
The objection I find hardest to dismiss comes from the civil-liberties direction rather than the industry one. Techdirt argues that provenance infrastructure creates an underappreciated surveillance surface linking identity to specific content, where a journalist signing footage before publication may unknowingly leave a server-side trace of it. Filmmakers and human-rights documenters were not the target, which is the problem with infrastructure: it does not know who it is for.
Even the enthusiasts are candid. The Transparency Coalition backed this bill, and its own explainer says it applies to only the largest AI systems with at least one million monthly visitors. When supporters concede the narrowness in the celebration post, that argument is over.
Which is why the follow-on matters. SB 1000 would delete the user threshold from the covered-provider definition and rename the detection tool a disclosure verification tool. The Senate concurred in the Assembly amendments 39-0 on August 27 — the last vote it needed — and, as I write on August 30, it is awaiting the Governor, unsigned. It carries an urgency clause, so if enacted it takes effect immediately. Check its status before you quote me.
What does this mean for you?
You are not going to read a statute this week. Here is the part that changes what you do.
Know what a detection tool answers. A compliant one tells you whether that company's system made the file. A clean result is not a verdict of "human" — it is "not theirs, as far as this tool can tell."
Run the check on your own files once.
exiftoolis free and the field is Digital Source Type. Check something you posted, then the copy the internet serves back. If the second is thinner, you have learned more in five minutes than any explainer will teach.Assume the mark is gone after the third hop. Screenshot, crop, resize, re-upload: every step strips metadata, and almost none of it is malicious. Provenance survives careful handling, not ordinary handling.
Do not read absence as innocence. Below a million monthly users nothing is required, text is outside the law, and two big generators already use incompatible schemes. "No label found" mostly means no label was found.
If you publish images, fill in the credit line too. In my test that plain-text field was the only survivor. The machine-readable one is more rigorous; the crude one is more durable.
If the threshold bothers you, this is a phone call, not a lawsuit. With no private right of action, enforcement is a choice made by public lawyers — and the threshold is a choice on a governor's desk.
The lesson, as I see it
I went into this expecting to write about a loophole and came out writing about maintenance.
The weak point is not the escape hatch, nor really the million-user line — a line is at least visible, and a bill to erase it has passed both houses. It is that a provenance regime is only as good as the least careful software in the chain, and no link in that chain is villainous. My image did not lose its label to a bad actor. It lost it to a thumbnail generator doing its job well.
Which is why the 2027 platform duty decides whether any of this works, and why "to the extent technically feasible" is not boilerplate but the actual policy. Feasibility is not a fact about the universe. It is a fact about what companies have already built — so the standard gets set by whoever moves first, then applied to everyone else.
Senator Josh Becker, who wrote the bill, said the goal isn't to tell people what to believe but to give them better information so they can decide for themselves. Right goal, delivery date attached. My vote? Judge this law in January 2027, not August 2026. Until the pipes are required to carry the signal, we are all admiring a label that gets thrown away somewhere between the server and your screen.
Go run the check on your own pictures. It takes one command, it will ruin your afternoon, and it is still cheaper than believing the wrong photograph.





