Microsoft Memo Casts AI Scraping as ‘the Largest Theft of Labor’

  • AI
  • September 18, 2026
  • 0 Comments

The documents had been redacted for months, and on September 17 a court lifted the black bars. Inside the copyright suit that The New York Times brought against OpenAI and Microsoft sat language that neither company would have chosen to make public.

Brent Hecht, a director of applied science at Microsoft, described what the company’s Copilot answer engine was doing to publishers in an internal presentation from January 2024. He called it a “doom loop.” Microsoft’s own measurements showed that Copilot had cut click-through traffic to The New York Times domain by as much as 93 percent when compared with a conventional Bing search.

One passage has drawn the most attention since the unsealing. “It is unusual for an end product to threaten the economic foundation of its key supplier,” the document reads, “and that is the situation we are in with the content supply chain of the LLM business.” Another internal line framed the practice of scraping for training data in blunter terms, calling it the largest theft of labor in history.

The statements matter because of what they contradict. OpenAI has built its defense in the litigation on the claim that training models on published work is fair use rather than theft. Microsoft, its largest backer and cloud partner, is a co-defendant. The unsealed files show employees at the two companies describing the practice in language that cuts against that position.

Satya Nadella, Microsoft’s chief executive, said in a deposition earlier this year that content sitting behind a paywall should be licensed if a company wants to use it. The remark was more measured than the internal presentation, but it pointed the same way: the people running the business understand that some content carries a price, even while the legal team argues it does not.

The Times filed its suit in late 2023, among the first major publishers to take the AI companies to court over the use of copyrighted work. The complaint argued that models trained on the paper’s journalism reproduce it without permission, and that the answer engines built on top of those models steer readers away from the source that paid for the reporting.

Internal documents have become the most damaging evidence in this class of case. What employees say in emails and slide decks is not constrained by the careful wording of court filings, and juries have tended to find the private language more believable than the public defense.

The 93 percent figure is the sharpest detail in the unsealed material. A product that answers a reader’s question directly, in a single paragraph, gives that reader little reason to click through to the newspaper that did the original reporting. If the traffic disappears, so does the subscription and advertising revenue that pays for the reporting in the first place.

The fair-use defense rests on a specific reading of copyright law, one that treats training a model as a transformative use of material rather than a reproduction of it. The companies argue that the models learn patterns, not passages, and that what they produce is new expression. The unsealed documents complicate that argument because they show employees thinking about the arrangement in commercial terms: a supplier being squeezed by a customer that depends on it.

The presentation’s framing of a “doom loop” is the clearest statement of the problem the publishers are trying to prove in court. Traffic drops when the answer engine replaces the search result, revenue drops with the traffic, investment in reporting drops with the revenue, and the quality of the material that trains the next model drops with the investment. The cycle, as Microsoft’s own staff described it, ends where it began.

Publishers have divided over how to respond. Some have signed licensing agreements with the AI companies, treating the model builders as a new kind of syndication customer. Others have gone to court, arguing that settling on the defendants’ terms would value their work too cheaply and set a precedent that chills the rest of the industry.

For Microsoft, the exposure is narrower than OpenAI’s in the eyes of some lawyers, because the company sells the infrastructure rather than training the models directly. But the unsealed documents place Microsoft’s own researchers in the room where the scraping was discussed, and that is the kind of detail that can widen a co-defendant’s liability.

The market has so far treated the litigation as a cost of doing business. Investors have continued to value the model builders as though the copyright question will be resolved through licensing deals rather than through adverse rulings. The unsealing does not settle that question, but it moves the companies closer to a day when a judge or jury reads their own words back to them.

The timing of the unsealing, in the middle of a legal fight that has already produced depositions from top executives, suggests the case is moving toward the stage where evidence becomes public rather than sealed. Each new document that emerges tends to shape the settlement math, and the settlement math is what most of these disputes ultimately resolve into.

Related Posts

  • September 23, 2026
  • 16 views
Anthropic and OpenEvidence to Give Free Medical AI to Poorer Countries

OpenEvidence began as a way for a doctor to ask a question and get an answer drawn from peer-reviewed research rather than a search engine. It is free for clinicians…

  • September 23, 2026
  • 20 views
Meta’s Muse Tops the Charts, Then Runs Into Amazon

Meta released Muse on Sept. 8 with a simple pitch: a personal AI agent that could book tickets, sort email and act across the web on a user’s behalf. The…