DOJ backs fair use for AI training as publishers challenge the evidence

The DOJ backs fair use for AI training, while publishers challenge its evidence and licensing arguments. With no ruling yet, the dispute turns on market substitution. Korean broadcasters’ lawsuit and Daum’s proposed payments show the stakes for media rights and revenue.

DOJ backs fair use for AI training as publishers challenge the evidence

The government’s intervention in the OpenAI litigation sharpens the dispute over news licensing and market substitution with implications for Korean broadcasters

The U.S. Justice Department has backed OpenAI and Microsoft’s fair-use position on AI training, drawing a detailed challenge from news publishers who say the government reached its conclusion before reviewing the evidence. The confrontation puts the economics of news production alongside the legal question of whether copying articles to build language models requires permission.

The department filed its statement of interest on September 1 in the consolidated copyright litigation before Judge Sidney H. Stein in the Southern District of New York. News publishers responded on September 28, asking the court to give the intervention no weight. Neither filing is a judgment. The court must still evaluate the conduct at issue and the statutory fair-use factors.

For Korean broadcasters, the dispute has a direct commercial relevance. KBS, MBC and SBS filed their own action against OpenAI in Seoul in February. Their claims concern the use of news for training and the potential diversion of audiences to AI answers. The American proceedings will not determine Korean law, but they show what evidence publishers may need when asserting harm to advertising, subscriptions and licensing markets.

DOJ position and publishers’ response. Source: ECF 2071. These are disputed arguments.

Table 1  Positions on training and fair use

Issue

DOJ position

Publishers’ response

Copyright Office analysis

Training

Highly transformative

Requires case-specific evidence

Consider copying and overall purpose

Outputs

Training defense does not automatically apply

Products may substitute for news

Competing expressive outputs matter

National security

Constraints may benefit foreign rivals

Policy cannot replace statutory factors

Report analyzes copyright factors

Licensing

Potential entry barrier

Other inputs already require payment

Availability is relevant with other conditions

Decision maker

Government submits a position

Court must decide on the record

Case-by-case judicial analysis

Evidence

Filed before summary-judgment briefing

Two years of discovery and 70+ witnesses

Public comments and case law


A government argument focused on training

The DOJ’s position concerns the copying of copyrighted works to train large language models. It characterizes that use as highly transformative and argues against a broad rule making AI training generally impermissible without licenses. The same defense does not necessarily extend to every output a model produces. Training, retrieval and the reproduction of protected expression in an answer remain distinct acts requiring careful analysis.

The publishers’ response identifies the government filing as ECF 1682 in case 25-md-3143. Their own response is ECF 2071. The litigation brings together claims from The New York Times, Daily News publishers, the Center for Investigative Reporting, The Intercept and Ziff Davis entities. Reporting on the summary-judgment motions describes approximately 10.8 million asserted articles. That figure measures the works in dispute, not a finding that each was infringed.

The government also presents licensing as a possible barrier to entry. It argues that burdens on domestic AI development could weaken U.S. national security and benefit foreign rivals, and that licensing requirements could entrench the largest developers. Publishers reject the link between those policy concerns and the legal treatment of the specific copying in this case.

Table 2  DOJ arguments as cited in the publishers’ response

DOJ brief location

Position as quoted in publishers’ response

Page 2

Domestic AI restrictions could affect national security

Page 4

Licensing barriers could entrench large developers

Page 4 footnote 13

Specialized content licensing agreements exist

Page 8

Fair use requires case-by-case analysis

Pages 12 and 17

Training characterized as highly transformative

Page 16 footnote 17

Criticism of the Copyright Office’s reasoning

Page 19

Opposition to broad liability requiring licenses for training


Publishers challenge the timing and factual basis

The September 28 response makes the evidentiary record its starting point. According to the publishers, the DOJ filed before the parties’ summary-judgment motions and before substantial portions of those filings became public. They say the department did not examine hundreds of documents, depositions from more than 70 fact witnesses or dozens of expert reports collected during two years of discovery.

That objection matters because fair use depends on the particular works, conduct and markets at issue. An argument that AI development benefits the public does not itself establish how a defendant obtained articles, what permissions applied or whether its products compete with the publisher’s services. The publishers ask the court to decide those questions on the record rather than substitute industrial policy for the statutory inquiry.

The response also disputes the treatment of the U.S. Copyright Office. It says the DOJ dismissed the office’s reasoning in footnote 17 instead of engaging with its analysis. Reporting about interagency consultation remains reporting, rather than a judicial finding about the filing process. The more consequential legal disagreement is visible in the documents themselves: how narrowly the relevant use should be defined and how the model’s ultimate commercial purpose should affect that analysis.

The Copyright Office’s May 2025 pre-publication report does not declare every instance of AI training unlawful. It examines copying at different stages and emphasizes the circumstances of each use. Its discussion of pirate-source copying, competing expressive outputs and reasonably available licenses identifies a combination of facts that would weigh against fair use. Those conditions should not be flattened into a general rule that all available licenses defeat the defense.

The office also cautions against treating dataset compilation or training as the entire purpose of an AI system. Separate acts of copying require separate consideration, but the overall use provides context. A model designed to generate expression substantially similar to protected works presents a different question from a genuinely different use of those works. The report rejects the proposition that training is inherently transformative simply because it is described as nonexpressive.

The White House’s national AI legislative framework, released on March 20, 2026, adds another layer. It states the administration’s belief that model training does not violate copyright while acknowledging contrary arguments and leaving resolution to the courts. It also recommends considering collective licensing arrangements without deciding when licenses are required. Supporting such arrangements is compatible with defending fair use in some circumstances; the dispute is whether the DOJ adequately addresses the evidence and the practical consequences in this case.

The publishers also cite an October 2023 FTC submission, made under the previous administration, about possible competition and consumer-protection issues associated with certain uses of protected expression. That document raises another regulatory question; it does not independently decide the pending copyright claims.

Selected public milestones. Sources: court filings, Copyright Office, White House and Korean reports.

Licensing costs and infrastructure spending

The publishers’ response challenges the claim that content costs would materially obstruct entry into AI. It cites reported testimony that Microsoft had spent more than $100 billion on its OpenAI partnership and reported projections of $175 billion in Microsoft capital expenditure in 2026. It also cites an estimate of $1.1 trillion in combined capital spending by Google, Amazon, Microsoft and Meta between 2023 and June 2026.

Those amounts cover different periods and categories. They should not be treated as a common annual AI-training budget, or used to calculate a meaningful licensing-cost ratio. The publishers deploy them to argue that chips, electricity and financing already impose large costs, and that content should not be singled out as an input that must be available without payment. Source

Different categories and periods. Sources: reports cited in ECF 2071 and the Daum media-day account.

The response contrasts those figures with the reported OpenAI–News Corp agreement valued at $250 million over five years. Reuters and Hearst also appear as examples of content licensing. Existing deals show that some transactions are feasible; they do not establish an industry-wide clearing price or demonstrate that every publisher and developer can negotiate on equal terms. Source

Ziff Davis chief executive Vivek Shah has separately argued that infrastructure spending is the more substantial barrier and that services such as Cloudflare, TollBit and Really Simple Licensing can reduce transaction friction. Music rights organizations such as ASCAP and BMI offer precedents for organized licensing, although news and audiovisual rights present different ownership and usage questions. The publishers’ filing likewise argues that large developers should attempt workable licensing mechanisms before treating them as impossible. Source

The supply problem behind the doom loop

The publishers cite internal technology-company material discussing the possibility that AI products substitute for the sources they depend on. Their response describes a Microsoft executive’s concern that an AI content strategy could weaken publishers’ ability to fund fresh journalism, eventually harming the models’ supply of current information. The filing is evidence of the publishers’ argument and its cited materials, not a corporate admission establishing liability. Source

The commercial mechanism is straightforward. A user may obtain enough information from an AI answer to avoid visiting the original site. If that reduces a publisher’s revenue and reporting capacity, fewer original stories may become available for future retrieval and training. The severity of each link must be established with evidence. A conceptual loop illustrates the risk; it cannot measure how much revenue any defendant caused a publisher to lose.

Conceptual supply risk advanced by publishers. Source: ECF 2071. Not a quantified causal finding.

The response invokes Prism News, which described scanning more than 10,000 sources, and reporting about Brown Brothers Media’s use of acquired publications for large volumes of AI-generated material. These examples support the publishers’ concern about competition for attention. Claims about their traffic, staffing or operating status remain attributed accounts, not findings by Judge Stein. The broader question is how an AI information service can sustain access to fresh reporting when its business model may reduce the incentive to produce it.

OpenAI disputes reproduction and market harm

OpenAI and Microsoft sought summary judgment on September 4. According to PPC Land’s account of the filings, the defense cites 24 verbatim outputs identified in 20 million ChatGPT conversations, equivalent to 0.00012 percent when divided by the conversation count. This is a defense-side statistic from a particular analysis, not the probability that any AI answer infringes copyright.

Selected figures from the parties and reporting. Counts do not establish an infringement rate.

The defendants also distinguish Bing search snippets from requests for full publisher content in browsing, and invoke implied permission for a period before publishers blocked the ChatGPT-User crawler. Whether a crawl was authorized depends on the applicable facts and legal requirements. A delay in blocking access should not be described as automatic consent to every subsequent training or output use.

The original article placed The New York Times’ 20.7 percent digital advertising growth in 2025. The cited account assigns that growth to the second quarter of 2026. The defense also points to the newspaper’s subscriber growth. Strong aggregate performance may inform the debate, but it does not by itself establish that particular conduct caused no harm to an existing or potential licensing market.

Publishers, for their part, distinguish literal reproduction from economic substitution. A service can answer a reader’s question without reproducing an article word for word. Whether that conduct affects a legally relevant market remains a matter for the court. The parties are therefore not simply disputing the same output count; they are contesting what kinds of use and market effect the fair-use analysis must capture.

The publishers cite Judge Vince Chhabria’s discussion in Kadrey v. Meta to argue that requiring permission need not end technological development. Meta nevertheless won summary judgment on the record in that 2025 case. The opinion’s warnings about market harm should be read alongside that outcome, rather than presented as a general holding against AI training.

A wider pattern of media litigation interventions

The source article places the AI filing alongside other DOJ interventions involving media companies. It describes the government’s support for Paramount Skydance’s demand for a $1.88 billion bond in litigation over its Warner Bros. Discovery acquisition, as reported on September 17. A September 21 settlement report describes additional U.S. production commitments of $1.5 billion over five years and arrangements concerning editorial governance at CBS and CNN.

Other examples in the source include a statement supporting Children’s Health Defense in a 2025 antitrust dispute involving news organizations, the March 2026 Live Nation–Ticketmaster settlement and closely timed DOJ and FCC action on Nexstar–Tegna. These cases supply policy context. They do not establish that the copyright filing was improper, or that consolidation determines whether training is fair use. Each proceeding has its own legal tests and record.

Media-related proceedings discussed in the source article. Their legal tests differ.

Korean broadcasters bring their own claims

KBS, MBC and SBS filed suit against OpenAI in the Seoul Central District Court on February 23, alleging copyright infringement and violations of unfair-competition law. Their complaint concerns news allegedly used without permission for model training and AI answers that they say divert traffic. The source article reports a combined damages demand of approximately KRW 400 million, about $295,000 using its reference exchange rate, together with requests concerning training data deletion. These are demands, not damages awarded.

The broadcasters’ counsel, Kim Tae-kyung, identifies market substitution as a central consideration. The source also cites Korea Broadcasters Association chairman Bang Moon-shin’s criticism of differing licensing treatment for Korean and overseas media, and his commitment at the association’s April meeting to respond to unauthorized use of broadcast content. Those remarks state the broadcasters’ position. The lawfulness of the uses remains disputed.

An American fair-use decision would not bind a Korean court. Korean copyright and unfair-competition claims must be evaluated under domestic law. The practical relevance is the evidence: which material entered which system, how users encountered it, what terms governed access and whether the resulting services affected markets for the original reporting.

Table 3  The U.S. and Korean proceedings

Item

U.S. consolidated litigation

Korean broadcasters’ action

Court

S.D.N.Y., 25-md-3143; Judge Sidney Stein

Seoul Central District Court

Plaintiffs

NYT, Daily News publishers, CIR, Intercept, Ziff Davis entities

KBS, MBC and SBS

Initial filing

NYT filed December 27, 2023

February 23, 2026

Works

Approximately 10.8 million asserted articles

News content; count not stated here

Claims

Copyright infringement and disputed market effects

Copyright and unfair competition

Relief sought

Damages and other remedies; disputed claims

About KRW 400m and data deletion requests

Status

Summary-judgment motions filed

Complaint filed; merits disputed

Government role

DOJ supports fair use for training

Administrative guidance is not a ruling


Financial pressure does not establish AI causation

The source cites Korea’s June 19 broadcast-market release and reporting of 2025 figures: total broadcast revenue of KRW 18.6495 trillion, down 0.8 percent; broadcast advertising of KRW 2.0134 trillion, down 12.3 percent; and terrestrial television advertising of KRW 693.6 billion, down 17 percent. The figures describe a sector already under pressure from shifting advertising and audience behavior.

They cannot identify how much of that decline came from AI answers. Establishing attributable harm would require evidence linking particular uses to lost visits, subscriptions, advertising revenue or licensing opportunities. Comparing annual advertising losses with a complaint’s damages demand also does not reveal the ultimate value of the rights at issue: the amounts serve different purposes.

2025 broadcast-market changes. Source: June 19, 2026 market release cited in Korean reporting.

The source describes a separate public project that spent KRW 20 billion to build 23,113 hours of rights-cleared broadcast training material, approximately 4.6 million records selected from more than two million hours of source programming. KBS, MBC, MBC Chungbuk and KT ENA material were reported among the inputs. A curated dataset can help demonstrate that licensed supply is possible, but its costs and terms do not automatically set prices for commercial access to broadcasters’ entire archives.

Korea also tests payment and licensing mechanisms

The Korean policy debate includes objections to proposals described as use first and compensate later. Creator groups argued that such an approach would weaken control over protected works, while government explanations distinguished news, newspapers and music with identifiable owners and existing markets. Administrative guidance released by the Ministry of Culture, Sports and Tourism and the Korea Copyright Commission in February addresses the familiar factors of purpose, nature, amount and market effect. Guidance informs the analysis; it does not decide a pending case.

Daum’s September 21 media-day proposals supply a more concrete commercial example. The source article incorrectly identified Daum as a Kakao-operated service and placed its AI-citation budget at KRW 20 billion. The Journalists Association of Korea’s September 22 account identifies Upstage’s May acquisition of the operator and an annual AI-citation budget of KRW 2 billion, approximately $1.47 million at the source article’s reference rate.

The proposed allocation would measure AI quotation use and distribute payments quarterly among participating outlets that consent to the use of their data. Daum also outlined data API packages for search, full-text access and training. Its illustrative revenue split is 70:30 after processing costs; the account should not be read as a universally executed contract. The report further describes a move toward fixed news fees and restrictions on external crawlers and AI agents.

These mechanisms address different activities. Payment for quotations in answers is not automatically a license for historical model training. The reported Amazon–New York Times agreement, worth up to $25 million annually, is a bilateral license rather than a shared allocation budget. The agreements should be compared by scope, duration, rights and calculation method, not by headline amounts alone.

Table 4  Payment proposals and content supply arrangements

Party

Mechanism

Reported scale or timing

Daum under Upstage

Proposed quarterly allocation by AI quotation use

KRW 2bn annually; September 21, 2026

Daum

Proposed data API packages

Illustrative 70:30 split after processing costs

Korean public dataset project

Rights-cleared audiovisual training data

KRW 20bn; 23,113 hours; about 4.6m records

Amazon–NYT

Bilateral AI content license

Reported annual ceiling of $25m

OpenAI–News Corp

Bilateral content agreement

Reported $250m over five years

Naver–KBS and Chosun Ilbo–Upstage

Individual partnerships

Terms not public in this article


A proposed shared budget and a bilateral license have different scopes. KRW conversions are indicative.

The source also cites Dong-A Ilbo strategist Kim Hyun-ji’s August 2025 warning that separate negotiations could weaken publishers’ bargaining power. Naver–KBS and Chosun Ilbo–Upstage arrangements illustrate a market developing through individual partnerships. Collective licensing might reduce friction, but participants would still need to settle ownership, permitted uses, reporting standards and allocation rules.

Table 5  Policy and commercial questions in the two markets

Issue

United States

South Korea

Policy

DOJ statement supports training fair use

Debate over use-first compensation proposals

Copyright authority

2025 report and DOJ offer different analyses

MCST and copyright commission provide guidance

Litigation

Consolidated publisher case; motions pending

KBS, MBC and SBS complaint

Payments

Individual licenses including News Corp

Daum proposes AI quotation allocation

Commercial question

Evidence of substitution and licensing feasibility

Rights scope, usage records and compensation


The next decision concerns evidence and usable contract terms

Judge Stein’s treatment of the summary-judgment motions is the next major legal milestone. A decision could resolve some issues, leave others for trial or narrow the disputed conduct. Denial of one motion would not by itself establish infringement across all 10.8 million articles. The government’s filing may influence argument in other cases, but it does not replace a ruling on the record.

Korean broadcasters face an additional rights problem beyond news text. Video, sound, subtitles, performers’ interests and metadata may have different owners and contractual restrictions. A pricing unit based on quoted text cannot simply be applied to an entire drama episode or archive. Contracts need to distinguish training from retrieval and output, specify the permitted material and provide usable records of access and payment.

The immediate task for media companies is to document the rights they control and the effects they can substantiate, while negotiating clear terms for authorized uses. A sustained decline in advertising is a reason to examine the business model; it is not proof of AI-caused loss. The dispute will become commercially actionable when the parties can identify what was used, under which rights, for which product and with what measurable consequences. Those details will determine whether new licensing revenue supports the next round of original reporting.