A Microsoft director privately called his own company's AI training practices "the largest theft of labor in human history." That line comes straight from an internal January 2023 memo, and it only became public on September 17, 2026, when a federal judge in Manhattan unsealed hundreds of pages of filings in The New York Times' copyright suit against Microsoft and OpenAI. The documents do not just allege wrongdoing from the outside. They show the companies' own staff describing the practice in terms a plaintiff's lawyer could not have written better.

What did the unsealed filings actually say?

The author of that line is Brent Hecht, Microsoft's director of applied science. In the memo, he wrote that the industry's wholesale scraping of copyrighted text amounted to "an astonishing theft of unprecedented proportions." A year later, in a January 2024 internal presentation, he returned to the subject with a term that has since become the story's shorthand: the "doom loop." His own words, quoted in the filing: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

RelatedJudge Approves Anthropic's Record $1.5B Copyright Settlement

Hecht wasn't describing a hypothetical. Microsoft's own internal data, cited in the same filing, shows that Copilot's AI-generated answers cut click-through rates to nytimes.com by as much as 93% compared with a traditional Bing search result. A reader who used to click through to the source article now gets a synthesized answer and never visits the site that reported it. Elsewhere in the documents, Hecht goes further still, conceding that OpenAI and Microsoft's fair use defense "would make a complete mockery" of the legal doctrine if it were allowed to stand.

The "doom loop" Microsoft's own memo describes Journalism gets scraped into AI training data, powers Copilot and ChatGPT answers, which divert reader clicks away from the original publisher, cutting the revenue that funds the next round of reporting. OriginalNYT journalism Scraped intoAI training data Copilot / ChatGPTanswers the question Reader neverclicks through (-93%) "THE DOOM LOOP" (Hecht, Jan 2024) Lost revenue funds less of the next story genztech.blog
Fig 1 Microsoft's own internal presentation described this cycle by name, months before the Times amended its suit to include it.

How big was the scraping, and did anyone try to hide it?

The filings put real figures on what had previously been argued in the abstract. One dataset cited in the case contained more than 91,692 copies of New York Times, Daily News and other investigative works. A separate training corpus, internally named Project Mango, held at least 160,903 unique works pulled from news publishers generally. The Times' broader claim, that OpenAI and Microsoft copied millions of its copyrighted articles in full, is the headline number in the suit itself.

More damaging than the scale is the conduct the documents describe around it. The filing alleges employees at OpenAI circumvented the Times' paywall and stripped copyright management information from articles before they went into training sets, exactly the kind of digital fingerprint the DMCA protects. One exchange quoted in the unsealed record has OpenAI researcher Nick Ryder telling company president Greg Brockman about a method for getting past the paywall. Brockman's reply, in full: "ah nice." It is the kind of casual line that reads very differently once it is evidence in federal court.

Why does Microsoft keep saying one thing internally and another in public?

In sworn testimony cited in the case, CEO Satya Nadella said that "anything that is paywalled should be licensed" before it is used to train a model, a standard Microsoft's own products appear not to have met with the Times' content. Brockman, for his part, is quoted describing generative AI products as posing an "existential threat" to publishers precisely because the outputs "are largely substitutive" for the original reporting, the legal standard that matters most under fair use analysis. Courts weighing a fair use defense look hard at whether a new use substitutes for the market of the original work. An OpenAI cofounder using that exact word, on the record, is not a coincidence a defense team wanted made public.

RelatedBig Tech's $1T AI lease bill isn't on the books

Microsoft / OpenAIGoogleAnthropicPerplexity
Stance on news contentLitigating; disputes it owes licensingSigned multi-year licensing deals with several publishersSettled and licensed with News Corp, othersRevenue-share "Comet" publisher program
Paywall handling (alleged)Circumvented, per unsealed filingsCrawls under robots.txt rules, disputed separatelyNot a party to this suitAccused of ignoring robots.txt by several outlets
Current legal exposureActive federal suit, trial-readiness decision expected 2027Multiple smaller suits, largely settledMostly resolved via licensingOngoing disputes, no major suit yet
  1. Dec 2023The New York Times sues OpenAI and Microsoft alleges unauthorized use of millions of articles to train GPT models
  2. Jan 2023 & Jan 2024Hecht's internal memo and "doom loop" presentation are written not yet public
  3. Sep 17, 2026Manhattan federal court unseals the filings the quotes above become public for the first time
  4. 2027Judge expected to rule on whether the case proceeds to trial fair use defense is squarely at stake

What does this mean for the stock?

Microsoft is not going to lose Copilot over one lawsuit, but the exposure here is bigger than a single plaintiff. The Times is the named party, but the same training corpora that pulled in 91,000-plus of its articles almost certainly pulled in comparable volumes from every other major outlet with a paywall. A ruling that rejects the fair use defense would not just cost damages against NYT content specifically, it would reopen the licensing question against the entire news industry Microsoft and OpenAI trained on without a deal. For investors, the signal to watch is not the headline damages figure, which historically settles for far less than initial claims. It is whether the court's language on the "doom loop" evidence narrows the fair use standard for AI training generally, since that ruling would apply well beyond MSFT and OpenAI to every model trained on scraped web text.

What to watch · 2026-2027
  • Discovery keeps surfacing internal admissions. Once a company's own staff calls a practice theft in writing, that language follows the case through every subsequent filing, deposition and appeal.
  • Other publishers watch the "doom loop" evidence closely. Any outlet with a paywall and a Copilot or ChatGPT citation problem now has a template for what discovery might turn up in its own dispute.
  • The fair use ruling, whichever way it goes, becomes the reference case for every other AI-training copyright suit still working through the courts.

Our take

What makes this different from the usual AI-copyright argument is that nobody outside the companies had to make the case. Microsoft's own applied-science director made it for them, in writing, twice, a year apart. Corporate memos calling something "the largest theft of labor in human history" do not normally survive to see daylight, and the fact this one did says more about how discovery works in federal litigation than it does about anyone's intentions to be candid. The "doom loop" framing is also the more useful phrase long term. It describes, in the company's own language, why publishers keep losing negotiating leverage against the tools built on their work: the traffic drop that funds fewer reporters shows up before any court ruling does.

Primary sources

Original analysis by GenZTech, based on unsealed Manhattan federal court filings and independent reporting. Source: TechCrunch