Newspaper publishers sue Microsoft and OpenAI over copyright
Nearly 400 newspaper publishers have filed a lawsuit against Microsoft and OpenAI in the U.S. District Court for the Southern District of New York. The publishers allege the companies unlawfully copied and used their copyrighted articles to train large language models without permission or compensation. They are seeking a court ruling declaring the unauthorized use as copyright infringement and an order to prevent further use of their content.

*this image is generated using AI for illustrative purposes only.
Nearly 400 newspaper publishers have filed a lawsuit against Microsoft and OpenAI, alleging the companies unlawfully copied and used their copyrighted articles to train the large language models behind their AI products without permission or compensation. The legal action, filed in the U.S. District Court for the Southern District of New York, claims the practice violates federal copyright law. The publishers argue that their journalism is protected by copyright and that utilizing it to train AI models necessitates authorization and payment.
Microsoft and OpenAI are partners in the development and deployment of AI systems designed to generate text-based responses. The plaintiffs are requesting a court ruling that the alleged unauthorized use constitutes copyright infringement. Additionally, they are seeking an injunction to prevent further use of their content by the technology firms.
This lawsuit contributes to a growing series of legal challenges initiated by publishers, authors, and various content creators. As courts deliberate on the application of copyright law to AI models trained on publicly available and copyrighted material, the outcome of this case could set a significant precedent.
Other technology companies are facing similar scrutiny regarding their AI training practices. In March, Grammarly encountered a lawsuit concerning its alleged use of an Expert Review AI tool. Julia Angwin, a contributing opinion editor at The New York Times, alleged that the tool utilized her name and the names of others without prior consent.
Furthermore, Anthropic, the AI entity responsible for the Claude chatbot, is currently facing legal action from music rights management company BMG. BMG asserts that Anthropic used lyrics from major artists to train its chatbot without securing the necessary authorization.
How might a ruling in this case influence the future cost structure and development timelines for training large language models?
If the court grants an injunction, what technical solutions or alternative data sourcing strategies could AI developers employ to continue model advancement?
Could this legal pressure accelerate the adoption of data licensing agreements between AI firms and content creators?






























