Internal Docs Show OpenAI and Microsoft Feared Web Harm
Internal emails and memos from OpenAI and Microsoft warned that their large-scale data scraping and AI models would damage the web, undercut publishers, and erode the very content supply chain that trains the models, even as the companies continued the practice for financial gain.
GGLOBAIPOLICY DESKSHARE
Internal emails and memos from OpenAI and Microsoft warned that their large-scale data scraping and AI models would damage the web…
Share this post
Short answer: Internal emails and memos from OpenAI and Microsoft warned that their large-scale data scraping and AI models would damage the web, undercut publishers, and erode the very content supply chain that trains the models, even as the companies continued the practice for financial gain.
internal docs reveal openai microsoft feared web harm
On September 18, 2026, The Verge reported that recently unsealed court filings in the New York Times lawsuit against OpenAI and Microsoft contain stark warnings from the companies themselves. The documents reveal that internal communications described the data-scraping practices used to train AI models as the largest ever taking of human labor and said the approach makes a joke of fair use.
The filings include remarks from Brent Hecht, Microsoft’s director of applied science, who warned that the companies were creating a self-reinforcing cycle that would damage the web. Microsoft later said those comments reflected only one employee’s personal view and did not represent the company’s position. Jordan Usdan, who leads data strategy and operations for Microsoft AI, characterized Hecht’s role as adversarial and academic, noting that he is employed to bring futuristic perspectives but does not speak for Microsoft on the effects of AI on creators.
Other executives also appear in the material. Satya Nadella acknowledged that chatbots have largely replaced traditional search, removing the incentive for users to visit original sources. An internal Microsoft memo noted that the AI content strategy had started a cycle that would hurt both model performance and the wider web, observing that it is rare for a product to undermine the economic base of its own suppliers, yet that is exactly what had happened with their large language model business and its content supply chain.
greg brockman motivated by financial returns from commercial ai
OpenAI cofounder Greg Brockman is described as being motivated by the prospect of enormous financial returns from commercial AI. An OpenAI representative admitted that the company had not made any effort to detect or remove paywalled material from its training data, despite later statements from Nadella suggesting that paywalled content should be licensed.
Internal discussions also showed that OpenAI recognized its models tended to memorize and regurgitate copyrighted text verbatim. Employees noted that preventing memorization was important to limit copyright violations, yet they acknowledged that GPT-4 had stored a vast amount of data and would therefore be highly prone to repeating passages. The filing cites multiple instances where ChatGPT output long strings taken directly from articles published by the New York Times, Mercury News, The Denver Post, LifeHacker and Eurogamer.
Microsoft acknowledged that the widespread scraping of online content was not what most creators intended, and that those creators receive no compensation for the use of their work. Jack Clark, OpenAI’s policy director, warned that the company was building systems that replace the labor of individuals who shape cultural discourse. Internal documents likened ChatGPT to a modern newsstand, and Nick Turley of OpenAI said that once a user obtains an answer from the chatbot there is little reason to click through to the original source.
microsoft admits scraping harms creators destroys ai supply chain
Microsoft further admitted that large language models effectively destroy their own supply chain because they often substitute for the very data used to train them. OpenAI’s own media and economic analysts linked the decline in referral traffic to news sites directly to AI-generated summaries such as Google’s AI Overviews, estimating that search referrals could have fallen by as much as 60 percent.
A Microsoft spokesperson later clarified that Satya Nadella’s testimony and the company’s official stance remain consistent, emphasizing that his comments addressed broad shifts in how people find and consume information and should not be conflated with the legal arguments about copyright that are still before the court.
Taken together, the unsealed material shows that both OpenAI and Microsoft were aware that their AI strategies risked severely harming publishers, the millions of people employed in the media industry, and ultimately the viability of their own products. Despite these warnings, the companies continued to pursue the approach, driven by the prospect of substantial financial gain. For developers and users of AI, the revelations underscore the importance of scrutinizing training data practices, considering the long-term health of the content ecosystem, and seeking ways to build models that do not undermine the creators whose work makes the technology possible.
Frequently asked questions
What internal warnings did OpenAI and Microsoft express about the impact of their AI data-scraping practices on the web?
Internal communications described the scraping as the largest ever taking of human labor, said it makes a joke of fair use, warned of a self-reinforcing cycle that would damage the web, and noted it hurts model performance and the wider web by undermining the economic base of its suppliers.
What did Brent Hecht, Microsoft’s director of applied science, warn about the companies’ AI strategy?
He warned that the companies were creating a self-reinforcing cycle that would damage the web, and Microsoft later said his comments reflected only one employee’s personal view and did not represent the company’s position.
How did OpenAI handle paywalled material in its training data according to the filings?
An OpenAI representative admitted that the company had not made any effort to detect or remove paywalled material from its training data, despite later statements from Nadella suggesting that such content should be licensed.
What estimate did internal analyses give for the decline in referral traffic to news sites due to AI-generated summaries?
OpenAI’s media and economic analysts linked the decline in referral traffic to news sites directly to AI-generated summaries such as Google’s AI Overviews, estimating that search referrals could have fallen by as much as 60 percent.
In mid-September 2026, researchers from Hacktron AI used Anthropic’s Claude tool to breach an OpenAI employee’s ChatGPT account, gaining access to private GitHub code after exploiting a misconfiguration in OpenAI’s Discourse forum.
California Governor Gavin Newsom issued an executive order on September 18, 2026, establishing a task force to recommend AI safety rules, including a mandatory kill switch for advanced systems, regular testing of that switch, third-party audits, and loss-of-control reporting, while federal AI legislation remains stalled.
The United States almost launched a military strike on a Chinese vessel after an AI-generated intelligence report, produced by a chatbot that hallucinated details, falsely claimed the ship was transporting components for a Chinese nuclear arms program, prompting preparations for an interception before the error was discovered.
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.