OpenAI and Microsoft internal docs reveal doom loop concerns
Internal documents show OpenAI and Microsoft warned that scraping vast web text to train LLMs creates a self-reinforcing "doom loop" that reduces traffic to source sites, undermines the data supply, and threatens both the open web and model performance.
Short answer: Internal documents show OpenAI and Microsoft warned that scraping vast web text to train LLMs creates a self-reinforcing "doom loop" that reduces traffic to source sites, undermines the data supply, and threatens both the open web and model performance.
internal warnings AI data scraping threat to open web
On September 18, 2026, newly unsealed court documents from the New York Times lawsuit against OpenAI and Microsoft showed that the companies’ own internal warnings described their data-scraping practices as a looming threat to the open web. The filings, which span 92 pages, contain statements from executives and engineers who acknowledged that harvesting vast amounts of online text to train large language models could set off a self-reinforcing cycle that harms both the models and the sources they rely on.
One of the most striking characterizations came from Microsoft’s Director of Applied Science, who described the data collection as the largest theft of labor in human history and argued that the companies’ legal defenses undermined the principle of fair use. Microsoft later clarified that those remarks reflected the individual’s personal viewpoint and did not represent an official corporate stance or legal analysis. A separate filing quoted a Microsoft general manager for data strategy and operations, who said the director’s role was to bring forward-looking, academic perspectives to the team and that his opinions should not be taken as the company’s position on AI’s impact on creators.
doom loop AI reducing website visits hurting content creators
The documents repeatedly refer to a “doom loop” in which AI-generated answers reduce the need for users to visit original websites, thereby weakening the economic base that supplies the training data. An internal Microsoft memo noted that the firm’s AI content strategy had already begun this loop, observing that it is rare for a final product to jeopardize the livelihood of its essential suppliers, yet that is exactly what had happened with its large language model business and its content supply chain.
Executives from both firms also acknowledged the financial incentives driving the effort. OpenAI’s co-founder was described as being motivated by the prospect of enormous financial returns from commercial AI products. Meanwhile, an OpenAI representative admitted that the company had not implemented any systematic effort to detect or remove paywalled material from its training datasets, despite later statements from Microsoft’s CEO suggesting that paywalled content should be licensed.
Internal discussions at OpenAI revealed awareness that its models frequently reproduce copyrighted text verbatim. Employees noted that while preventing memorization was considered important to limit copyright violations, GPT-4 had retained a substantial amount of data, making it highly likely to regurgitate passages from source articles. The filing cited multiple instances where ChatGPT outputted long excerpts taken directly from pieces published by the New York Times, the Mercury News, The Denver Post, LifeHacker and Eurogamer.
creators not compensated for AI reuse of their work
Microsoft’s own assessments echoed the concern that creators rarely expect their work to be repurposed in this manner and receive no compensation for such use. OpenAI’s policy director warned that the technology was effectively substituting for the labor of individuals who shape cultural discourse, describing the chatbot as a modern newsstand. An OpenAI engineer later remarked that once a user obtains an answer from the model, there is little incentive to click through to the original source.
The documents also highlight the damage to the models’ own supply chain. Microsoft acknowledged that large language models can become substitutes for the very data they were trained on, thereby undermining the foundation of their own performance. OpenAI’s media and economic analysts linked the rollout of AI-powered search summaries, such as Google’s AI Overviews, to a noticeable decline in referral traffic for publishers like the New York Times, estimating that search-based visits could have fallen by as much as 60 percent.
A Microsoft spokesperson sought to frame the executives’ testimony as consistent with the company’s broader observations about shifting information-seeking habits, arguing that those remarks should not be conflated with legal conclusions about copyright that are still before the court. Nonetheless, the unsealed material makes clear that both OpenAI and Microsoft were aware that their AI strategies risked damaging the publishing ecosystem, jeopardizing the livelihoods of millions of people who create online content, and ultimately threatening the sustainability of their own products, yet they proceeded in pursuit of substantial financial gain.
For developers and users of AI systems, the revelations underscore the importance of scrutinizing how training data is sourced and considering the broader effects on content creators. When building or deploying models that rely on web-scale text, it is prudent to assess whether the approach inadvertently undermines the very sites that supply the information, to explore licensing or compensation mechanisms, and to monitor traffic impacts on source publishers. Staying informed about ongoing legal debates and emerging best practices for data governance can help mitigate the risk of contributing to a cycle that harms both the open web and the AI systems that depend on it.
Frequently asked questions
What did the internal documents reveal about the companies' data-scraping practices?
The internal documents showed that executives and engineers warned that harvesting vast amounts of online text to train large language models could create a self-reinforcing cycle that harms both the models and the sources they rely on, describing the practice as a looming threat to the open web.
What is the "doom loop" described in the filings?
The doom loop refers to a cycle where AI-generated answers reduce users’ need to visit original websites, weakening the economic base that supplies the training data, which in turn undermines the models’ performance and threatens the sustainability of both the AI products and the content creators.
Did the companies acknowledge financial motivations and lack of efforts to filter paywalled content?
Yes. OpenAI’s co-founder was said to be motivated by the prospect of enormous financial returns from commercial AI products, and an OpenAI representative admitted the company had not implemented any systematic effort to detect or remove paywalled material from its training datasets, despite later statements suggesting such content should be licensed.
Source: The Verge
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.