
Background
Pursuant to 28 U.S.C. § 517, the U.S. government filed a Statement of Interest of the United States in the multidistrict litigation before the U.S. District Court for the Southern District of New York concerning OpenAI’s use of copyrighted works to train large language models. This marks the federal government’s first direct intervention in a series of copyright cases involving artificial intelligence training. The government has urged the court to regard model training based on copyrighted text—as distinguished from the use of AI-generated outputs—as a highly transformative fair use under existing law.
Prior to this development, Shira Perlmutter, Register of Copyrights and Director of the U.S. Copyright Office, released a pre-publication version of the report Copyright and Artificial Intelligence, Part 3: Generative AI Training in May 2025. The Office had previously issued separate reports addressing digital replicas and the copyrightability of works. As noted in our previous client alert, the report concluded that whether the fair use doctrine applies to generative AI training must be determined based on the facts of each case. The outcome depends on such factors as the works used, the sources of the training materials, the purpose of the model, the model’s outputs and the relevant licensing markets.
Shortly thereafter, the Trump administration attempted to remove Perlmutter from office, prompting litigation over the supervisory authority of the Copyright Office and whether the President has the power to remove the Register of Copyrights, who is appointed and may be removed by the Librarian of Congress. On June 30, 2026, the U.S. Supreme Court denied the government’s application to stay a ruling of the U.S. Court of Appeals for the District of Columbia Circuit. That ruling allowed Perlmutter to remain in office while the litigation proceeded. It also preserved the possibility that the U.S. Copyright Office could formally issue the final version of its Part 3 report, as discussed in our previous legal briefing on the decision and its implications.
Against this background, together with the administration’s publication of America’s AI Action Plan last year, the position taken by the executive branch in this litigation is clear: following the Copyright Office’s more cautious and fact-specific analysis, the administration has adopted a position supporting AI training in the present case. This also demonstrates that AI training has become not only a copyright law issue, but also a policy priority at the government level.
The legal brief also places fair use within a broader national policy framework. The government links the development of artificial intelligence—and, more specifically, large language models—to national security, economic competitiveness, scientific progress and U.S. leadership in emerging technologies. It warns that judicially imposed licensing obligations for training activities could slow domestic AI development, strengthen foreign competitors and concentrate the market in the hands of companies capable of bearing the costs of large-scale licensing.
For corporate clients, the significance of this filing lies in its translation of policy considerations into litigation arguments. The government does not leave questions relating to AI training within a broad and undefined factual inquiry. Instead, it advocates a practical rule for adjudication: courts should distinguish between a model’s internal training processes and its externally generated outputs. In the absence of a separate act involving the substitutional use of protected expression, the mere use of copyrighted text to train a large language model should not, by itself, be found to constitute infringement.
The Government’s Arguments
The government’s arguments are based on the well-established fair use doctrine. However, the manner in which the brief applies that doctrine would substantially narrow the scope of copyright claims directed at the training stage. As the OpenAI–New York Times litigation moves toward the summary judgment stage, the following four points may be the most significant:
· Training Activities and Generated Outputs: The government argues that fair use must be assessed separately for each specific use. Internal or intermediate copying undertaken to train a model should therefore not be treated in the same manner as public-facing outputs that may reproduce protected expression. This distinction is critical because plaintiffs frequently rely on allegedly infringing outputs, summaries or memorized reproductions of text to challenge the training process itself. The brief urges the court to address any output-related issues through remedies directed specifically at those outputs.
· Large Language Model Training Is Highly Transformative: The brief describes training as a process in which text is used to identify linguistic patterns, associations and predictive signals, rather than to replace the expressive purpose of the underlying works. From this perspective, model training is different in nature from reading, displaying, selling or republishing the copyrighted works themselves.
· Commercial Purpose Is Not Determinative: The government acknowledges that many large language model products are commercial in nature. It nevertheless argues that commerciality should be given less weight where the challenged use is transformative and does not disclose protected expression to the public. This reasoning is intended to keep the first fair use factor focused on the purpose of the training itself, rather than on the business model of the AI developer.
· Harm Must Involve Substitution, Not General Competition: With respect to market harm under the fourth fair use factor, the brief rejects the view that AI-generated works on similar subject matter, or broader competition resulting from AI-generated content, constitutes cognizable copyright harm in itself. Instead, the government argues that copyright-related market harm requires substantial similarity, substitution for protected expression, or impairment of a recognized derivative market directly connected to the works at issue.
In addition, the government does not reject voluntary licensing models or rule out the possibility of future legislative solutions. It argues, however, that when adjudicating cases under existing law, courts should not convert unsettled policy questions into a mandatory licensing regime applicable to all training scenarios.
Differences from the U.S. Copyright Office’s Pre-Publication Part 3 Report
The most important contextual consideration is that the two documents emerged within the policy environment of the same administration, yet differ substantially in their central orientation. Shortly before the Trump administration attempted to remove Shira Perlmutter, the Copyright Office released the pre-publication Part 3 report in May 2025. The report concluded that the fair use analysis for generative AI training is highly dependent on the facts of each case and does not support a categorical conclusion.
The Part 3 report recognizes that fair use may protect certain training uses, particularly non-commercial research or analytical uses that do not result in the reproduction of protected expression. At the same time, the report warns that copying expressive works from pirated sources to generate unrestricted outputs capable of competing in the market is unlikely to qualify as fair use where licensing channels were reasonably available. The report treats the source of the training materials, the purpose of the model, safeguards governing outputs and developing licensing markets as central factual considerations in the fair use analysis.
By contrast, the government’s Statement of Interest asks the court hearing the OpenAI litigation to characterize training as an internal, transformative use and to treat outputs, the acquisition of training materials and licensing issues as separate matters. This represents a significant departure from the Part 3 report. Rather than considering differences among various training scenarios, the government seeks to establish a broad rule that the use of copyrighted text for training alone should not be regarded as infringement.
This shift is particularly pronounced with respect to licensing. The Copyright Office recognizes that voluntary licensing markets are still developing and suggests that collective licensing arrangements could be considered if market gaps persist. The government’s brief, however, warns that judicially imposed licensing obligations would weaken U.S. competitiveness in artificial intelligence and entrench large incumbent companies. It thereby treats licensing primarily as a legislative or commercial issue, rather than a requirement arising from the judicial application of fair use.
The two documents also differ significantly in their treatment of market harm. The Part 3 report includes potential licensing markets and market-competing outputs within the scope of the fair use analysis. By contrast, the government’s brief rejects a broad theory of “market dilution” and argues that competition created by artificial intelligence does not constitute cognizable market harm under the fourth factor unless the output is substantially similar to, or substitutes for, protected expression.
Implications and Key Takeaways
The government’s filing is not legally binding on the court, nor does it resolve the unsettled fair use disputes pending in the numerous AI copyright lawsuits. Nevertheless, it is still likely to influence legal briefing in the OpenAI–New York Times litigation and related disputes. On the one hand, it provides defendants with an official federal policy position that they may cite. On the other hand, it gives copyright holders advance notice of the defense arguments to which they will need to respond.
· For AI developers and enterprise users: The filing strengthens the defense that copying at the training stage may constitute fair use where the model does not disclose protected expression. This does not amount to absolute immunity. Claims based on facts involving materials obtained from pirated sources, memorized reproduction by models, substantially similar outputs or improper use of copyrighted content may still succeed depending on the specific circumstances. Courts may also reject the arguments presented in the brief and decline to recognize a fair use defense for AI training altogether.
· In the short term: Stakeholders should preserve records concerning the sources of training data, model safeguards, licensing history, output controls and all verified instances of reproduction or substitution. These facts, rather than abstract debate over the advantages and disadvantages of artificial intelligence, are more likely to determine the direction of the next stage of fair use litigation.


Follow us