A group of news publishers has filed suit against OpenAI and Microsoft, alleging that the companies used published content without authorization to train their AI models, according to MediaPost. The case is the latest in a wave of copyright litigation that has followed the industry’s adoption of AI, and it carries the potential to reshape how AI companies pay for the material they learn from.
The suit is not the first, but it is one of the largest by number of plaintiffs. Publishers from across the industry have joined, according to people familiar with the filing, arguing that their articles — reporting that cost money, time and professional risk to produce — were scraped in bulk and used to build systems that now compete with them for readers and advertisers.
The central legal question is familiar: does training an AI on copyrighted text constitute fair use, or does it require a license? The publishers say the use is commercial, massive and direct, and that the outputs of the models reproduce and paraphrase their reporting in ways that substitute for the original work.
OpenAI and Microsoft have defended the practice with the fair-use argument that has anchored the AI industry’s legal position: models learn patterns, not copies, and the public benefits of the technology outweigh the private claims of content owners. Courts have not yet settled the question, and the cases are moving through different jurisdictions with different standards.
The lawsuit follows The New York Times’s suit against OpenAI, filed in December 2023, which remains the defining case in the field. Since then, authors, visual artists and record labels have filed their own actions, and the courts have begun producing rulings that cut in both directions. The publishers’ case adds another layer to an already crowded docket.
The stakes for the AI companies are existential in scale. If the courts rule that training requires licensing, the cost of building future models rises sharply, and the industry’s economics change at the foundation. If the ruling favors fair use, the publishers lose their most valuable claim — and with it, the bargaining power over the platforms that now distribute their work.
The stakes for publishers are existential in a different way. The industry’s traffic has been falling as AI answers displace search referrals, and advertising revenue has followed. The suits are partly a legal claim and partly a business strategy: a way to force the AI companies to pay for content, either through court orders or through settlements that establish licensing as the norm.
The licensing wave has already begun outside the courtroom. OpenAI has signed content deals with major publishers, including News Corp, Axel Springer and the Associated Press, paying for access to archives and ongoing use. The new suit suggests that publishers outside those deals are not willing to wait for the terms to be set one by one.
The defense’s strongest argument is technical. Training data is processed statistically, and models do not store copies of articles in any meaningful sense; they store probabilities about language. But the plaintiffs point to outputs that reproduce passages nearly verbatim, and courts have shown interest in those examples. The line between learning and copying is the battleground.
The case also raises questions about the supply chain of AI training. The publishers’ suit targets both the model maker and the platform that distributes it, and it names Microsoft in part because of its investment in and partnership with OpenAI. The structure of the industry — one company building models, another selling them at scale — means liability questions attach to both.
For Microsoft, the suit arrives at an awkward moment. The company has been making its own deals with publishers and has positioned itself as a friend to the news industry through licensing and distribution partnerships. A suit that treats it as a co-defendant in unauthorized training undercuts that positioning.
The outcomes will take years to resolve. Copyright cases of this scale move slowly, and appeals are all but certain regardless of the initial rulings. In the meantime, both sides are negotiating in the shadow of the litigation, and the settlements signed so far are likely to be followed by more.
The deeper question is whether the industry can find a sustainable arrangement. Publishers want payment and control; AI companies want data and flexibility; readers want access to both good journalism and capable AI. The courts will set the legal boundaries, but the commercial settlement — who pays, how much, and for what — will be worked out in deals, not verdicts.
For the publishers, the immediate goal is simpler: a seat at the table. The suits, whatever their legal merits, have already changed the terms of the negotiation. AI companies that once treated news content as free input now sign licenses, and publishers that once had little bargaining power now have a claim. The fight over training data is really a fight over the future price of facts.
The case will be watched closely by every industry that produces creative work — books, music, images, journalism. Whatever the courts decide about news articles will set the pattern for the rest. The publishers’ suit is one front in a war that will define who owns the material that machines learn from, and who gets paid when they use it.


