Internal Microsoft communications from December 2023, unsealed as part of The New York Times' copyright lawsuit against OpenAI and Microsoft, highlight significant executive concerns about the economic repercussions of AI content scraping. Brent Hecht, Microsoft's director of applied science, predicted in an internal presentation that the company's Copilot answer engine could create "doom loops" that harm both AI model performance and the internet's overall health.
Hecht described this dynamic as one where AI extracts value from written content only to have its own output threaten the economic viability of its sources. He characterized this as "the largest theft of labor in human history" in a January 2023 memo. Microsoft's internal data, as presented in the filing, indicated that Copilot's answer engine caused a drastic reduction in click-through rates for The New York Times, with figures ranging from 87% to 93% lower than standard Bing searches. Similar impacts were observed on Ziff Davis domains, which include publications like Eurogamer and IGN, where click-through rates dropped between 51% and 94%.
The court filing also revealed internal discussions at OpenAI. A researcher named Nick Ryder reportedly informed OpenAI president Greg Brockman about a method to bypass The New York Times' paywall for scraping purposes, to which Brockman responded with apparent approval. Internal OpenAI documents suggest Brockman believed large language models (LLMs) are "very good at any news task."
While the common defense for AI training practices has been fair use, the unsealed documents present a challenge to this argument. The filing also notes that Microsoft CEO Satya Nadella stated under oath during a deposition that he would have "required OpenAI to retrain its models" had he known that paywalled information was being used for training.