DOJ Backs Fair Use for LLM Training — But Carves Out Model Outputs
The US Justice Department filed a Statement of Interest on September 1 in the OpenAI copyright litigation, arguing that training LLMs on copyrighted works is 'exceedingly transformative' fair use. It expressly reserves that output-stage uses may not be transformative.
The US Department of Justice filed a Statement of Interest of the United States on September 1, 2026 in the copyright infringement litigation against OpenAI1. Its argument: training large language models (LLMs) on copyrighted works is “exceedingly transformative” and constitutes fair use1. TechCrunch reported it as the US government siding with OpenAI on the training-data question2.
Reading this as “the government declared AI training legal” would be inaccurate. The filing addresses the training stage only, and it explicitly reserves that output-stage uses may not be transformative1.
A filing, not a ruling
Start with what this document is. A Statement of Interest lets the Justice Department inform a court of the United States’ interest in pending litigation, under 28 U.S.C. §5171. The statute carries no time limitation and does not require the court’s leave1.
In other words, the government stated its position in writing. It does not bind the court. The filing itself notes that the fair-use inquiry hinges on the specific facts and uses at issue in each case1. A footnote adds that the government does not contend any of the activities alleged or described in the litigation were authorized or consented to by the government1.
The filing says it refers only to the New York Times and OpenAI for simplicity, but that its legal arguments apply similarly to all parties in the litigation and related cases, including book authors and publishers1. The dispute is not treated as confined to one case.
The transformativeness argument
The government’s reasoning runs as follows. An LLM uses a copyrighted work not to duplicate its expressive content, but as part of a process to learn and act on statistical patterns in written text — vocabulary, syntax, and knowledge1. The purpose of the copying (building an intelligent, interactive model) differs in kind from the purpose of the copied work (using language to entertain or educate a reading audience)1.
Citing a 2025 decision that called copying to train an LLM “transformative—spectacularly so,” the filing concludes that “the use of copies to train LLMs is extraordinarily transformative”1. That OpenAI uses its LLMs commercially “does not move the needle,” it says1. On the fourth factor — the effect on the potential market — the filing concludes that it “heavily favors fair use”1.
The government frames its interest around AI industry competitiveness and national security1. Citing executive orders from January 2025 and June 2026, it states that sustaining America’s global AI dominance is US policy, and argues that “rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered”1.
The licensing-fee argument
The filing also makes a competition argument. An erroneous fair-use ruling would hamper competition in the LLM market, because only the largest technology companies might have the capital to pay licensing fees1. Such fees would disproportionately benefit legacy media outlets given the sheer volume of their written publications, and the filing states it “is not in the public’s interest for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers”1.
A caveat follows here too. In a footnote, the government takes no position on whether a licensing regime would be financially or logistically feasible1. The same footnote observes that regardless of whether training constitutes fair use, both mainstream and independent publishers can and do enter licensing agreements giving developers specialized access to content1. Training legality and the existence of a licensing market are treated as separate questions.
The reservation attached to outputs
The most easily overlooked part of the filing is the line it draws between the training stage and the output stage.
The government states that LLM development proceeds in stages and that each stage may present distinct questions of copyright law, then says it focuses on the training stage1. It then writes: “To be sure, at the output (rather than training) stage, certain uses may not be transformative if the LLM reconstructs and disseminates an original copyrighted work”1.
So the government concedes that fair use may not hold where a generated output reconstructs and distributes a copyrighted work. Its argument is that because a use-by-use analysis is required, developing a tool that understands patterns and makes predictions must be distinguished from how people might use that tool to generate particular outputs1. It is not defending outputs in order to defend training. On remedies, it similarly argues that a tiny sliver of anomalous reconstructive outputs would not support massive liability for output uses generally1.
That distinction lines up with other pending litigation. When music publishers sued Anthropic and two founders in August, the acquisition, training, and output of lyrics were each separately at issue. Judging stage by stage also echoes how Japan’s Agency for Cultural Affairs guidelines separate the training stage from the generation and use stage.
What changes, and what does not
The filing does not immediately change the legal position of anyone using generative AI at work, because it is not a ruling.
What shifts is the negotiating and argumentative environment. With the US government’s position on training data now committed to a court filing, future arguments on this point will be made with reference to it. Both rights holders and developers can be expected to factor that position into licensing negotiations.
The operational picture moves less. As the government’s own reservation makes clear, the residual risk sits on the output side. Checking whether text or images produced by a generative model reconstruct an existing work remains necessary regardless of how the training question resolves. For anyone writing internal usage rules, designing that output-review process is still the part that requires hands-on work. In software, projects are settling this for themselves in parallel — Debian passed a general resolution on the responsible use of generative AI.
Sources
- Statement of Interest of the United States - The original 20-page filing submitted by the US Department of Justice on September 1, 2026 in In re: OpenAI, Inc. Copyright Infringement Litigation (25-md-3143 (SHS) (OTW), US District Court for the Southern District of New York), via the CourtListener RECAP archive
- US government sides with OpenAI on issue of training LLMs on copyrighted material - TechCrunch (September 2, 2026)
Was this article helpful?
Thank you!
Received. Thank you!