English

LawsuitsMicrosoftCopilot

Microsoft Claims Massive Reproduction of Copyrighted Works via Copilot is Extremely Rare

This article is a translation. Read the Japanese original

Microsoft has submitted the results of an analysis of 8.2 million Copilot chat logs as part of its copyright infringement lawsuits with publishers and authors.

This data consists of logs screened by experts for keywords related to news sites. According to Microsoft, fewer than 1% of all logs contained common descriptions of 16 words or more.

Specifically, 59,545 cases showed commonalities with news content. Meanwhile, in data related to the authors' lawsuits, only 24 responses had matches of 30 words or more. Out of 212 books evaluated, matches were found in only 10.

Microsoft argues that these figures support the claim that using copyrighted works for AI training constitutes "fair use." The company asserts that even if some text is reproduced, it does not undermine the transformative purpose of training large language models.

Currently, the publishers allege that Microsoft and OpenAI are building direct competing products using copyrighted works. Microsoft is seeking a summary judgment from the court to resolve the matter early.


Source: Microsoft says virtually nobody was grabbing NYT articles through its chatbot (The Verge AI, 2026-09-05)