The first real rulings in the AI training copyright lawsuits have landed, and the score surprises both sides. Courts in 2025 held that training on legally acquired books can qualify as fair use, refused to extend that shield to training built on pirated libraries, and aimed their sharpest scrutiny at AI output that reproduces actual articles, photos, and songs. The New York Times, book authors, Getty Images, and the major record labels are all still in court, and the answer will define what copying means for a generation of software.

The split that matters to you is simpler. Training is a courtroom fight measured in years. Output is your fight, and it started the day your work appeared somewhere it shouldn't. An AI tool or content farm that regurgitates your article, image, or track is infringing now, and the DMCA machinery that has always removed copies removes these too.

The Core Question: Is Training on Copyrighted Works Fair Use?

Training requires copying. Every book, article, image, and song fed into a model is first duplicated onto the developer's servers, and reproduction is the core exclusive right under 17 U.S.C. § 106. Without a license, the whole defense runs through fair use, the four-factor balance in 17 U.S.C. § 107.

Two factors have swallowed the others in AI disputes. The first is transformativeness: does the new use serve a purpose different from the original, or compete with it? The Supreme Court's 2023 Warhol ruling sharpened that question, holding that a use sharing the original's purpose and market weighs against fair use. The second is market harm. A model that emits a cheap substitute for your photograph, article, or novel does damage no search index ever did.

The industry's favorite precedent is the Google Books case, where scanning millions of books to power a searchable index was fair use. But Google displayed snippets, never the books; nothing it showed could substitute for the original. Generative models can. That difference, what comes out the other end, is why the training rulings hedge and why output claims keep surviving.

The Cases That Matter, and What Courts Have Actually Ruled

News: The New York Times v. OpenAI and Microsoft. Filed in December 2023 in Manhattan federal court, this is the flagship text case. Its most damaging exhibits are prompts that coaxed the chatbot into reproducing Times articles nearly word for word. In early 2025 the court let the core claims proceed past dismissal into the evidence phase. Meanwhile, the market answered on its own: several major news organizations have signed paid licensing deals with OpenAI, the training-license market forming in real time.

Books: Authors Guild v. Anthropic and Kadrey v. Meta. In June 2025, a federal judge in San Francisco held that training on legally acquired books is fair use, but that amassing hundreds of thousands of books from pirate shadow libraries was not, sending that issue toward trial. Days earlier, the same judge had applied the same reasoning to a parallel Anthropic case brought by comic artists. Meta's case ended differently the very next day: summary judgment for Meta, because the authors presented no real proof of market dilution, a failure of evidence, the judge stressed, not a blessing of training. By fall 2025, Anthropic had settled the authors' class claims for a reported sum near $1.5 billion, paired with a framework for licensing works going forward.

Legal research: Thomson Reuters v. Ross Intelligence. The February 2025 Third Circuit ruling most commentary skips: training a competing legal-research tool on Westlaw headnotes was not fair use. Ross copied the expressive core of a product to build a substitute for that product. A general-purpose chatbot is an easier fair-use story than a model built to replace the very works it ingested, and the tension between this holding and the Ninth Circuit's book rulings is the kind of split that reaches the Supreme Court.

Images: Getty Images v. Stability AI. Filed in early 2023 in US and UK courts, with a centerpiece exhibit few forget: generated images bearing distorted copies of the Getty watermark, proof the model memorized specific files. The UK case went to trial in 2025; the US case continues in federal court. For enforcement, the process photographers use on any stolen image works on AI clones too.

Music: the major labels v. Suno and Udio. Filed in June 2024, with exhibits of outputs that mimic specific recordings down to the vocals. The litigation has pushed parts of the AI-music industry toward licensing talks and settlement. Individual artists need not wait: infringing outputs come down through the same enforcement track musicians have always used.

Training Copies vs. Model Output: The Line Every Ruling Draws

Every decision so far draws one line: inside the model versus out of it.

Training copies are internal and intermediate. Where they were lawfully acquired and the plaintiff cannot show concrete market harm, courts have leaned toward fair use, with explicit warnings that the holding is narrow and record-dependent.

Output is where the lawsuits bite. A model that can be prompted to emit your article, your watermarked photo, or a soundalike track is handing users your expression. Verbatim reproductions anchor the Times complaint; the watermark images anchor Getty's. When output substantially similar to your work appears anywhere, that is ordinary infringement, independent of how the model was built or how the training fight ends.

The training holdings are hedged and headed for appeal. The output theories are the claims that keep surviving, and the ones you can enforce without waiting for any court.

Why DMCA Takedowns Can't Reach Training, but Cover What Models Publish

The DMCA, 17 U.S.C. § 512, is built for hosting services: companies that store content at users' direction and keep safe harbor by removing material on notice. A training corpus is not that. The copies sit on the AI developer's own servers; no user uploaded your novel to those disks, and no notice can make anyone "take down" model weights. Enforcement against training runs through licensing demands and litigation, not takedown notices.

Output is a different story. The moment an AI-generated article, image, or song is posted, to a blog, a marketplace, a social platform, an app, it sits on someone's hosting stack, and § 512 works exactly as it always has. Send a properly formed takedown notice to the host, and the ordinary machinery takes over: expeditious removal, or the host's safe-harbor protection is at risk.

Two details matter. It makes no difference that the copy was machine-generated; infringement turns on the copying, not on the copier's own rights. And when an operator hides behind domain privacy, a § 512(h) subpoena can identify them through their service provider.

Your Output-Side Enforcement Playbook

Detect. You cannot enforce what you never see. Set alerts for distinctive sentences from your best-performing pieces, run reverse image searches on flagship photos, and consider canary strings, invented phrases planted in originals that surface only when someone copies you. At volume, ProtectionPro earns its keep, scanning continuously and pushing matches into a removal pipeline.

Triage. Before filing, read the output through a fair-use lens. A post quoting a paragraph to criticize your article is probably protected. A rewritten near-copy published as someone else's work is not. Rights holders must consider fair use before sending takedowns, the dancing-baby Lenz litigation made that a duty, and careless mass-filing invites counter-notices and § 512(f) misrepresentation claims, when not to file is required reading.

Enforce. Start with finding the host, because the host controls the copy. File per-URL notices, then climb the ladder: platform report, a Google delisting request, registrar contact, and repeat-infringer strikes against accounts running AI content farms. The full sequence, including what to do when a host ignores you, is laid out in the guide to getting stolen content off the web.

Watch for impersonation. When output clones your face or voice rather than your text, copyright may be the wrong tool. Publicity rights and platform impersonation policies take over, and deepfake removal runs on its own track.

Build the Evidence File Before You Need It

Whichever way the appeals end, your leverage is records.

Ownership. Keep originals with creation dates, working drafts, camera files with intact EXIF metadata, and archived publication pages. Timestamp evidence closes the gap between knowing you wrote something first and proving it; documenting ownership is the same discipline whether the copy came from a scraper or a chatbot.

Registration. A takedown needs no registration, but a US lawsuit does, registration is the statutory prerequisite under 17 U.S.C. § 411. Register before publication or within three months after it, and you preserve statutory damages and attorney's fees under 17 U.S.C. § 412, the difference between a nuisance settlement and a credible threat. Whether registration is worth the fee now has a clear answer for anything with licensing value, and the authors' enforcement playbook treats pre-publication registration as routine.

Market records. The Meta ruling turned on the authors' failure to prove market harm. Rights holders with documented licensing income, stock libraries, wire services, syndicated writers, can prove it. Track your licensing history and every refusal; that file is the raw material of the market-harm analysis courts decide on.

Where This Is Headed: Opt-Outs, Licensing Markets, and the Appeals Pipeline

In the US. The answer is arriving case by case. Elsewhere, statute moves faster. EU copyright law lets rights holders opt out of commercial text-and-data mining with machine-readable reservations, and the EU's AI Act began requiring general-purpose AI providers in 2025 to publish summaries of what their models were trained on, transparency US litigants are still excavating through discovery. If your hosting or audience sits in Europe, those rules may drive your outcome more than US fair use, so know how takedown rights travel across borders before you file.

A licensing market is forming around training data. News organizations have signed paid deals, image libraries have followed, and the Anthropic settlement pairs payment with permission going forward. Every deal signed makes the next "we copied first and negotiated never" defense look worse on market harm, which is why courts keep asking who sought permission first.

Practical stance while appeals and follow-on trials climb through the Ninth Circuit, the Third Circuit, and possibly the Supreme Court: publish terms reserving AI-training rights on your site, use the crawler opt-outs the major labs honor, and log every refusal. None of that makes training unlawful by itself. All of it strengthens your hand in negotiation and in court.

Frequently Asked Questions About AI Training Copyright Cases

Can I send a DMCA takedown against an AI company for training on my work?

No. Section 512 notices target copies a hosting service stores at users' direction; training copies sit on the developer's own servers, and model weights are not removable content. Your remedies against training are licensing demands and litigation. Takedowns apply the moment training output is published somewhere, that is the copy you can reach.

Has any court actually ruled that AI training is fair use?

Yes, narrowly. In June 2025, federal courts held that training on legally acquired books is fair use, while holding that acquiring books from pirate libraries is not, and granted Meta summary judgment because the authors failed to prove market harm. The Third Circuit went the other way for a tool trained to compete with the works it copied. Appeals could reshape all of it.

If an AI tool reproduces my article or image verbatim, is that actionable?

Output substantially similar to your protected work can infringe regardless of how the model was built, the theory behind the strongest exhibits in the Times and Getty cases. Where the output is published, takedown notices work like any other copy. Suing the AI company itself is a separate, harder path, and what AI-generated material can own never excuses copying yours.

Do I need to register my copyright before sending a takedown?

No. A valid notice needs your work, the infringing URLs, and the required statements, not a certificate. Registration matters one step up: it is required to sue in the US, and timely registration preserves statutory damages and attorney's fees. Takedown today; register the works with real licensing value this week.

Does blocking AI crawlers with robots.txt stop my content from being used in training?

It stops the crawlers of companies that honor them, and the major labs honor their opt-outs today. But it is respected by policy, not court order: it does not reach datasets already built, copies gathered through third-party scrapers, or services that ignore it. Use it, document it, and pair it with site terms reserving training rights.

Your Next Five Moves

  1. Register your highest-value works now. Timely registration preserves statutory damages and attorney's fees, leverage that shapes settlements long before any trial.
  2. Turn on monitoring: alerts for signature phrases, reverse image checks, and an automated service if you publish at volume.
  3. Clear today's infringements today, locate the host, send per-URL notices, escalate to search delisting and registrar contact, and stack strikes against repeat offenders.
  4. Publish AI-training reservation terms and block the crawlers you can; keep a dated log of every opt-out and refusal.
  5. Escalate deliberately. When a licensing offer, settlement approach, or counter-notice lands, that is the moment for deciding when to bring in a copyright lawyer, before you respond, not after.