It was a pleasure to burn.It was a special pleasure to see things eaten, to see things blackened and changed. With the brass nozzle in his fists, with this great python spitting its venomous kerosene upon the world, the blood pounded in his head, and his hands were the hands of some amazing conductor playing all the symphonies of blazing and burning to bring down the tatters and charcoal ruins of history. With his symbolic helmet numbered 451 on his stolid head, and his eyes all orange flame with the thought of what came next, he flicked the igniter and the house jumped up in a gorging fire that burned the evening sky red and yellow and black. He strode in a swarm of fireflies. He wanted above all, like the old joke, to shove a marshmallow on a stick in the furnace, while the flapping pigeon-winged books died on the porch and lawn of the house. While the books went up in sparkling whirls and blew away on a wind turned dark with burning.
—— Ray Bradbury (Fahrenheit 451)
In many cases nowadays when we are searching for our materials online, we sometimes do this with a grey area—as long as you have any browser downloader or some random torrents, you could have really lots of free stuff, but most of the time, these actions already consist of illegal acts (such as violations of the DMCA).
Of course, whether you can avoid legal liability depends on whether the copyright holder chooses to pursue the matter further. As for whether the government will pursue the matter on behalf of the copyright holder, that depends on who you are.
Aaron Swartz - he helped develop the web feed format RSS and also was the co-owner of Reddit. From 2010 to the start of 2011, he downloaded 4.8 million papers from paid academic database JSTOR through MIT's intranet (almost 80% of the full data that time).
Aaron Swartz, a Data Crusader and Now, a Cause - The New York Times
FAQ | Report to the President: MIT and the Prosecution of Aaron Swartz.
He didn't publicly release these files.
So JSTOR decided to take the hardware from him and reached a civil settlement with him (The Register), which meant this case would no longer be pursued. But federal prosecutors did not stop their move; their charges escalated from 4 to 13, including two counts of wire fraud and 11 counts of violating the Computer Fraud and Abuse Act (CFAA), carrying a maximum potential sentence of 35 years and a fine of $1 million.
"Stealing is stealing whether you use a computer command or a crowbar, and whether you take documents, data, or dollars."
—Carmen Ortiz (federal prosecutor at the time)
He became a scapegoat of the new information age. What broke him down wasn't the from public opinion : any plea agreement had to include an actual prison sentence, or else the case would go to trial (Rolling Stone). On January 11, 2013, less than three months before his trial was set to begin, Swartz died in his Brooklyn apartment at the age of 26. After his suicide, the prosecution dropped their charges.
His actions were indeed illegal in the law. But when we look back at how large language models were originally trained nowadays—and what they were trained on—the whole thing is truly ridiculous.
The Greatest Theft Of All Time?
On September 17, 2026, a batch of internal documents originally marked as confidential were unsealed in The New York Times' copyright lawsuit against OpenAI and Microsoft (TechCrunch). Microsoft's director of Applied Science, Brent Hecht, wrote in a January 2023 internal memo that he called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
OK, let's back to the Swartz.
He downloaded 4.8 million papers, didn't even release them, and handed over the hard drives. He faced 13 federal felonies. In Meta's cases, this company downloaded at least 81.7 TB of data via BitTorrent through Anna’s Archive, including 35.7 TB from Z-Library and LibGen; prior to that, it had also downloaded 80.6 TB from LibGen (Ars Technica).
Anthropic also downloaded hundreds of thousands of books from two pirated book repositories, LibGen and PiLiMi. The judge ruled that using pirated content to establish a “central library” did not constitute fair use, and the case ultimately settled for $1.5 billion: involving approximately 480,000 works, at about $3,000 per work. (JURIST、Authors Guild).
Now you know the differences?
An individual who downloads millions of papers faces criminal charges and decades in prison; a company that downloads hundreds of terabytes of data faces a civil lawsuit. In its worst-case? Nothing more than a “settlement expense” on the financial statements.
No one ends up get caught.
Same here for the images, people try to put some unidentified noise into the image to fool AI in 2023 to protect it from AI crawling. Such as Glaze and Nightshade: the prior one will jam your style from AI's vision, and Nightshade will let the AI model inject "posions" so that they will understand the image in an absurdly wrong way (University of Cambridge).
In fact, these adversarial protections, just like a grocery store, compete with NASA on building rockets. These programs can't directly fight with a trillion-dollar industry.
And the result was that adversarial protections could not protect artists at all.
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
It could be easily broken down by some simple bypass like LightShed (the counter program of Nightshade). It can achieve 99.98% accuracy in identifying Nightshade-poisoned images, can “wash away” perturbations, and can even transfer what it has learned to other tools like Mist and MetaCloak(USENIX、MIT Technology Review).
You will never know whether these companies were able to delete such protection from you. Anyone can easily conduct targeted training to circumvent a particular defense, while content creators themselves are unable to effectively protect their own work. There is virtually no legal oversight of the training process, and conglomerates' lobbyists are keep their persuasion and messing around everywhere.
We should realize that what we are facing is a more general knowledge-information inequality.
Until now, no U.S. appellate court has ruled on whether “training AI with copyrighted works constitutes fair use”; in 2025, two federal judges heard similar cases and reached conflicting conclusions (gHacks). The only clear-cut issue is the line drawn against building databases with pirated content—paying settlements.
As for the data itself, names like Books3 and LibGen appearing in one lawsuit after another; even with current datasets, it is difficult to provide evidence proving they are truly compliant. Training data is a black box for court and us, and the key is held by the these conglomerates.
Oh, What did you say? Watermarks?
Article 50 of the EU AI Act requires that AI embed the label “I am AI” in its output. Starting August 2, 2026, providers of generative AI must apply machine-readable labels to text, images, audio, and video so that they can be detected; deepfakes and AI-generated text involving the public interest must also be accompanied by labels visible to the naked eye(Paul, Weiss).
It should protect like all of us from... what? These unethical training...? Right?
The EU is likely trying to use watermarks to distinguish between “AI-generated content” and “non-AI content" and to protect us in between the process. But if you consider this: once AI content is watermarked, any stolen content instantly becomes “AI-generated.”
Watermark only answered one question: Is it machine-generated or not? It does not answer another question: whose works does it make it learn to write to draw like this? It is labelled as "AI," not the millions being devoured. It is an utter steal: creators didn't gain extra benefits and protections from this policy. On the contrary, all the origins and the processes have been taken away and occupied by a sole watermark.
Ben Thompson reached a similar conclusion from a different perspective in a post on Stratechery. He said he has a deep philosophical objection to watermarking: AI (at least for now) is a tool in human hands, and forcing it to apply watermarks is like requiring a ballpoint pen to declare itself the author on a piece of paper; the EU is effectively ceding to AI the last thing humans have left—creativity—along with the credit for it (Write Things Down、Anthropic's Watermarking).
A 26-year-old programmer downloaded academic papers and was hounded by the state apparatus to the corner; a group of conglomerates with billions in net worth swallowed up nearly all of humanity's knowledge, news articles, and artworks—only to face a few lawsuits, pay a few settlements, and receive a “statement of interest” from the Department of Justice.
They can do this not only because the law hasn't kept up but also because the numbers add up. A $1.5 billion settlement, when poured into an industry that invests hundreds of billions of dollars annually, is nothing more than an operating cost. In the government's sights, it's an add-on of “national competitiveness.”
Does the number really add up? Training materials may be relatively 'free,' but the computing power, chips, data centers, and electricity—every single item costs.
And the answer to where this money comes from is even more bizarre than one might imagine: a large portion of it comes from these companies investing in, buying from, and lending to one another.
In the end, we're not only asking, "What did they steal?"
Instead, "Who's paying all of this?"
No comments:
Post a Comment