Sekėjai

Ieškoti šiame dienoraštyje

2026 m. rugsėjo 22 d., antradienis

Both OpenAI and Anthropic Are Run By Thieves Who Accomplished the Greatest Theft of Human Work Results in History --- Filing Details OpenAI Book Use --- Documents show employees weighed cost of buying titles to train ChatGPT. Stop these people now. Don’t allow them to steal our trade secrets too.

 

Unsealed court documents from mid-September 2026 have brought internal communications from OpenAI, Microsoft, and Anthropic into the public eye. These records, released as part of high-profile copyright lawsuits filed by news publishers and authors, show that employees and executives privately debated the ethics, legal risks, and costs of the data used to train artificial intelligence models.

Key Revelations from the Court Filings

           Book Training Debates: Unsealed briefs from the Authors Guild v. OpenAI case show that OpenAI employees discussed the high costs of buying book titles legally to train early iterations of ChatGPT. Internal messages revealed discussions about using unauthorized sources, such as the book-pirating repository Library Genesis, with one 2019 internal note stating, "We trained GPT-3 on pirated stuff! No sharing that!".

 

     "Theft of Labor" Internal Concerns: A legal brief unsealed in The New York Times copyright lawsuit showed that Dr. Brent Hecht, Microsoft’s director of applied science, privately warned in 2023 that mass data scraping could be viewed by the public as "an astonishing theft of unprecedented proportions" and the “largest theft of labor in human history”.

 

           Industry Displacement: Nick Turley, the head of ChatGPT at OpenAI, wrote in internal notes that AI posed an "existential threat" to publishers and that their products were "largely substitutive". Similarly, Jack Clark—former OpenAI policy director and co-founder of Anthropic—warned in 2020 that their systems would substitute for human authors on platforms like Amazon and make people unemployed.

 

     Anthropic's Legal Rulings: Legal actions have targeted Anthropic as well. While courts have ruled that legally purchasing print books to scan and digitize them for AI training falls under "fair use," Anthropic recently paid a $1.5 billion settlement to authors over claims that it used millions of pirated books from ebook piracy sites.

 

The Tech Industry Defense

Microsoft and OpenAI have continuously maintained that scanning public internet data and published works to train AI models is legally protected under the "fair use" doctrine of U.S. copyright law. They argue that the training process does not copy the work to resell it, but rather transforms the data to teach a machine how human language and logic function. Microsoft also stated that the "theft of labor" comments represented the perspective of an individual research academic and did not reflect the official stance of the corporation.

The legal battles remain ongoing in federal court as judges weigh summary judgments to determine whether the foundational data collection practices of generative AI companies constitute copyright infringement or fair use.

 

We are concerned about how these scraping practices might impact our own intellectual property or trade secrets.


“Authors including John Grisham, David Baldacci, Jodi Picoult and Jonathan Franzen sued OpenAI and Microsoft three years ago for copyright infringement.

 

A court filing unsealed on Thursday details internal messages and testimony laying out how OpenAI employees -- including executives -- talked about the company's use of pirated books to train an early ChatGPT model. The filing was made in support of the authors' request that a federal judge rule in their favor ahead of a trial.

 

The tech companies argued their actions constituted "fair use," which allows for some copyright material to be used without explicit permission, and that their use of the material was transformative.

 

Starting in 2019, OpenAI began using books from file-sharing site Library Genesis -- LibGen for short -- to help train its large language model. The site, which provides free access to books and scholarly articles, faced accusations of pirating copyright material.

 

OpenAI's then-general counsel David Lansky was among those who recommended free book downloads from LibGen as a possible data source, according to the filing.

 

Tom Brown, then a top GPT-3 engineer, and Ben Mann, another member of the technical staff, described LibGen as "sketchy AF," according to messages and testimony cited in the filing. A federal court ordered LibGen to shut down in 2015 and academic publisher Elsevier received a $15 million judgment against the site in 2017.

 

OpenAI staffers discussed removing references to LibGen from papers that could later be public, according to the filing.

 

In one instance, Dario Amodei, then a senior researcher at OpenAI who is now chief executive of Anthropic, asked in Slack whether it was "sketchy to call our corpuses 'Books1' and 'Books2' and not say what they are, particularly when in fact they are Libgen (which is a slightly sketchy source)."

 

Separately, Mann wrote that a description of the data sets was "deliberately vague since it's libgen."

 

A number of those involved in OpenAI's early model training now work at Anthropic.

 

OpenAI said employees who used LibGen to create the earlier ChatGPT models are no longer at the company, and the LibGen data set wasn't used to power the current ChatGPT models. An Anthropic spokeswoman declined to comment. Convergent Research, where Lansky now works, didn't respond to a request for comment.

 

OpenAI eventually deleted the pirated data sets. Bob McGrew, then vice president of research, wrote in a June 2022 Slack channel, "now is the right time to excise Libgen from our systems and storage," given how much OpenAI was in the news.

 

Removing LibGen would prevent researchers from being able to reproduce earlier GPT-3 or GPT-3.5 results, he said, adding, "but would be very valuable for legal reasons."

 

OpenAI removed the data sets from its library.

 

The company priced out the potential cost of buying large quantities of books on multiple occasions. Greg Brockman, OpenAI president, wrote to OpenAI co-founder Ilya Sutskever in late 2022 that they could purchase lots of books, "but they are expensive per token so we haven't prioritized." (Tokens are the basic measurement unit for AI use.)

 

Varun Shetty, OpenAI vice president of media partnerships, testified in the lawsuit that he wasn't aware of the company purchasing books at scale and scanning them for training purposes, according to the filing.

 

Safe Superintelligence, Sutskever's current AI lab, didn't respond to a request for comment.

 

Wall Street Journal parent News Corp has a content deal with OpenAI.” [1]

 

1. Filing Details OpenAI Book Use --- Documents show employees weighed cost of buying titles to train ChatGPT. Korn, Melissa; Bruell, Alexandra.  Wall Street Journal, Eastern edition; New York, N.Y.. 22 Sep 2026: B4. 

Komentarų nėra: