Skip to content
The Big Picture September 3, 2026

Washington Just Argued in Court That Your Content Is Free Training Data

The Justice Department told a federal judge on September 1 that training on copyrighted text is fair use, and that licensing fees would mostly enrich the biggest archives. Yours isn't one.

By The State of AI Marketing newsroom
Share
Editorial illustration for: Washington Just Argued in Court That Your Content Is Free Training Data
Credit: JAC Growth Marketing

On September 1 the Justice Department filed a 20-page brief in the consolidated OpenAI copyright case telling a federal judge that training an AI model on copyrighted writing is fair use. If you publish content at a small B2B software company, that filing is the U.S. government arguing in open court that nobody owes you anything for the words you put on the internet. It goes further. Any market where AI companies do pay for text, the brief says, should pay the largest archives first.

A statement of interest is what the government files to tell a court where it stands without joining the case. The filing went to Judge Sidney Stein in the Southern District of New York. He is hearing the suits from The New York Times, The Intercept, Tribune and Ziff Davis together. It’s the first time Washington has taken a public side in this wave of AI copyright suits, and it sides with OpenAI.

The legal headline is the transformation argument:

“In sum, the use of copies to train LLMs is extraordinarily transformative.”

LLMs are large language models, the systems behind ChatGPT and Claude. The brief’s point is that they read text to learn language rather than to resell it. That part was expected. The passage worth your attention sits on page 4, long before that, where the government says what it thinks a licensing market would do:

“An erroneous fair use ruling would hamper competition in the market for LLMs, because only the largest technology companies might have the capital necessary to pay licensing fees. And such licensing fees would disproportionately benefit legacy media outlets due to the sheer volume of their written publications.”

Read that as a marketer rather than a lawyer. The government’s position goes past “training is free” and lands on “paying for training would be bad policy, and if anyone does get paid it should be whoever has the most words.”

The pivot everyone made has a missing half

For two years the standing advice has been to stop chasing rankings and start getting cited inside AI answers. That advice is fine. Underneath it sat a second belief, mostly unspoken. Being in the training set and in the answer would eventually convert into a bargaining position, a licensing check, a seat at a table. This brief argues that bargaining position out of existence, in the government’s own voice.

Matt Topic, the Loevy & Loevy partner representing The Intercept, told his own newsroom what he thinks the filing does:

“If the administration’s position was accepted, it would result in an unprecedented, uncompensated transfer of IP rights from news organizations to tech companies.”

The word doing the work is uncompensated. Topic represents a newsroom that runs on member donations, not one with a 170-year archive, and a transfer of this kind runs downhill. The publishers with the least to bargain with lose first, because they had the least to lose.

The Times took the same line in blunter terms. Spokesperson Graham James said the administration was siding “with a handful of trillion-dollar AI companies” and that “AI companies simply need to pay fairly for the content that makes their products possible.”

Why the volume argument cuts against you twice

Licensing markets for text price by the pound, because volume is the only thing that’s cheap to measure across millions of documents. The Times has an archive going back to the 1850s. A 25-person B2B SaaS has maybe 400 blog posts and a docs site. In a market where AI companies pay for text, you’re a rounding error. In a market where they don’t, you’re free.

Both outcomes end the same way: the check doesn’t arrive. The only difference is whether somebody else’s does.

Having argued that licensing would subsidize old media, the brief claims models will “help level the playing field between mainstream and independent publishers.” Its example: a writer with limited resources could use a model to generate an image instead of hiring a photographer. The field gets leveled by making the small publisher’s costs lower, not their work worth more. Same argument, same page.

Publishers have already been rebuilding their sites so agents can read them cheaply. A few have gone the other way and started poisoning the crawlers they used to block. Pew’s classifier found that commercial .com domains carry AI authorship at ten times the .edu rate. The supply of machine-readable text was already growing fast. A brief arguing it should also be free is policy catching up to an economics that had already settled.

The strongest argument against us is in a footnote

Footnote 13 of the same brief undercuts the flat reading:

“Notably, regardless of whether LLM model training on text articles constitutes fair use, both mainstream and independent publishers could enter (and have entered) into licensing agreements to provide developers with specialized access to real-time, pay-walled, proprietary, and other content and information.”

The government makes that distinction on purpose. Training on published text is one thing. Selling live, gated, proprietary access is another, and the second market survives a fair use ruling untouched. The brief also says outright that it takes no position on whether a licensing regime would even be workable.

So the licensing market isn’t dead. It moved. What it pays for now is the thing a crawler can’t already take: your live pricing feed, your customer data, your benchmarks, your product telemetry. Almost no B2B content program has any of that. A blog doesn’t qualify, and neither does a gated report that was crawlable last week.

The verdict

Judge Stein decides this, not the Justice Department, and a statement of interest carries no binding weight. But the government has now put its position in writing, and the position is that your published words are raw material.

Keep doing the work that gets you cited. Visibility inside an AI answer is still the closest thing to free distribution you have. Just stop budgeting for the second half of that story. Your content is a marketing asset and it’s very unlikely to become a revenue asset. If you want something an AI company would pay for, it has to be something they can’t read for free. That isn’t what you’ve been publishing.

Quoted in this story

  • Matt Topic, Partner and counsel for The Intercept, Loevy & Loevy (source)
  • Graham James, Spokesperson, The New York Times (source)

Want your perspective in coverage like this? Get quoted.

Sources

This story is part of our running coverage: the full picture →

Get Net Effect.

The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.

More from The Big Picture