Did Meta Just Win the AI Copyright War? Why the Headlines Are Wrong

An Insights briefing from Muzy & Meraris LLP

By Muzamil Naeem

8/26/20267 min read

an image of an infinite sign on a blue background
an image of an infinite sign on a blue background

When a US federal court ruled that Meta's use of copyrighted books to train its Llama artificial-intelligence model was "fair use," the headlines wrote themselves: Meta wins. AI training is legal. The authors lose. For the technology industry, it read like vindication — confirmation that the vast appetite of generative AI for training data had the law's blessing.

That reading is wrong, or at least dangerously incomplete. Read the actual ruling, and a very different picture emerges: a court that handed Meta a narrow, record-specific victory while making clear it found much of Meta's argument unpersuasive — and, in the process, wrote a roadmap for the next wave of plaintiffs to win. For any business building, deploying or investing in AI, the lesson is not "training data is free." It is closer to the opposite.

This briefing explains what the court actually decided, why the victory is far narrower than reported, the separate and still-live question of how the data was obtained, and the practical steps every business using AI should take. A note at the outset: this concerns US copyright law and ongoing litigation, outside our firm's jurisdiction of admission, and is general information only, not legal advice. The law here is unsettled and moving quickly; specific advice from qualified counsel should be taken on any AI or data question.

The background: authors versus the machine

The case, brought in the Northern District of California, was filed by a group of authors — including well-known novelists — who alleged that Meta had trained its Llama large language model on their copyrighted books without permission. Their complaint had two distinct strands, and keeping them separate is essential to understanding everything that follows.

The first strand concerned the use: that training an AI on their books was itself copyright infringement. The second concerned the acquisition: that Meta had obtained those books by downloading them from "shadow libraries" — repositories of pirated works — via torrenting, which involves not only downloading but also simultaneously distributing copies to others.

In mid-2025, the court granted Meta summary judgment on the fair use defence in respect of the training. On its face, a win. But the reasoning is everything.

Fair use, briefly

Under US law, "fair use" is a defence to copyright infringement that permits limited unlicensed use of protected works in certain circumstances. Courts weigh four factors: the purpose and character of the use (including whether it is "transformative"); the nature of the copyrighted work; the amount and substantiality of what was used; and — critically — the effect on the potential market for the original.

The fourth factor, market harm, is generally regarded as the most important. And it is the factor on which this entire case turned.

Why Meta's "win" is far narrower than it looks

Here is what the headlines omitted, and what any business relying on this ruling must understand.

Meta won on the record, not on the principle. The court found the training use "transformative" — the books were used to teach a model to generate language, not simply republished. But it did not declare, as a matter of law, that training AI on copyrighted works is always fair use. On the contrary, the judge made clear that he found several of Meta's fair-use arguments unpersuasive, and that the outcome rested on the specific, limited evidentiary record before him.

The authors lost because they failed to prove harm — not because they were wrong. The court's central point was that the plaintiffs had not produced sufficient evidence that Meta's model harmed the market for their works — that its outputs would substitute for their books, or that Meta's unlicensed use had diminished their ability to license their works for AI training. Market harm was the decisive factor, and the plaintiffs simply had not built the record to prove it.

The court effectively wrote the next plaintiff's playbook. In explaining why these authors lost, the judge illuminated exactly how a future plaintiff might win: by proving genuine market harm — that an AI's outputs act as market substitutes for the originals, or that the defendant bypassed a functioning licensing market. The ruling reads less like a door closing on authors than like a set of directions for walking through it. The court also rejected the "ridiculous" suggestion that adverse copyright rulings would halt AI development — pointedly noting that these products are expected to generate enormous sums, and implying that such value can bear the cost of licensing.

The honest summary: this was a technical, fact-specific win for Meta, not a landmark declaration that AI training is lawful. Anyone treating it as settled permission to train on copyrighted material is building on sand.

The separate question that hasn't gone away: how the data was obtained

Even setting aside the use of the works, there is a second, distinct liability that this ruling did not resolve — and it may prove more dangerous.

The allegation that Meta obtained the books by downloading them from pirated "shadow libraries," and distributing them in the process of torrenting, concerns not what Meta did with the works but how it got them. That is a different legal question, and those claims remain live.

The distinction matters enormously, and a parallel case makes the stakes concrete. In closely related litigation against another AI developer, a court took the view that while training may be fair use, storing and using pirated copies is a different matter — and that case settled for a reported US$1.5 billion, with an estimated payout of around US$3,000 per work. In other words: even where the use is defensible, the unlawful acquisition of the training data can carry a colossal, separate liability.

The lesson is stark. Transformative use is not a licence to obtain the inputs unlawfully. A business may have a strong fair-use argument for what it built, and still face ruinous liability for how it sourced the material.

The next wave is built to win

Perhaps the clearest sign that Meta's victory is not the end of the story is what has followed it. Major publishers have since launched fresh litigation against Meta over Llama — and, crucially, these newer cases appear designed around the roadmap the court provided.

Where the original authors failed to prove market harm, the new plaintiffs are reportedly marshalling exactly that evidence: alleging that the AI's outputs produce full-length articles, journal papers, replacement textbook chapters and study guides that substitute for their works and directly displace their sales; that the company sourced enormous volumes of material from pirate sites while masking its tracks; and that it abandoned legitimate licensing negotiations. They are, in effect, litigating to the exact specification the earlier ruling laid out.

And there is a further, powerful fact now in play: a licensing market demonstrably exists. Meta itself has signed content-licensing agreements with major publishers and news organisations. In fair-use analysis, the existence of such deals is significant — it shows there is a market to license training data, that participants have priced and negotiated entry to it, and that the company chose to license some works while bypassing others. That undercuts a fair-use defence for the works it did not license, and strengthens the market-harm argument considerably.

This is not a settled area of law. It is an actively contested frontier, and the balance may yet shift decisively toward creators.

What businesses should actually take from this

For any company that builds, deploys, invests in, or acquires AI systems, the practical lessons are clear and urgent.

Do not treat "AI training is fair use" as settled — because it isn't. A narrow, record-specific ruling is not a green light. The law is unresolved, the next wave of cases is stronger, and the direction of travel is toward requiring licences. Building a business model on the assumption that scraping copyrighted data is lawful is a serious risk.

Licence your training data. This is the single most important takeaway. The transformativeness of what you build will not protect you if you cannot account for how you obtained the inputs — and the emergence of a functioning licensing market makes "we couldn't have licensed it" an increasingly hollow defence. Licensed data is defensible data.

How you acquire data is a distinct liability — treat it as such. Even a strong fair-use position does not cure unlawful acquisition. Training on pirated material, or from sources you cannot verify, exposes you to separate and potentially enormous claims. Provenance is not a footnote; it is central.

Diligence the data behind any AI you buy or build on. For acquirers and deployers, the question is not only "does this model work?" but "what was it trained on, and was that lawful?" As earlier guidance in this area has emphasised, liability travels with the model. A model trained on tainted data carries that exposure to whoever owns or uses it.

"Everyone else did it too" is not a defence. The scale of the industry does not immunise it. With generative AI projected to generate extraordinary revenues, courts are increasingly unmoved by the argument that requiring licences would be too burdensome — the value at stake cuts the other way.

A concluding observation

The story that "Meta won the AI copyright war" is a textbook example of why the headline is not the holding. What actually happened was narrower, more conditional, and far more instructive: a court gave Meta a fact-specific victory, told the world it found much of Meta's reasoning unconvincing, and effectively coached the next generation of plaintiffs on how to prevail. Meanwhile, the separate question of how AI companies obtain their training data — through licences or through piracy — remains a live and potentially ruinous liability, as a US$1.5 billion settlement elsewhere has already shown. For businesses, the comfortable reading of this ruling is the dangerous one. The prudent course is the opposite of complacency: assume the law is unsettled and tightening, licence your training data, verify its provenance, and treat the sourcing of information as the legal act it has become. In the AI economy, the value is in the data — and increasingly, so is the liability.

Published this briefing for general awareness. It concerns US copyright law and ongoing litigation, outside the firm's jurisdiction of admission; it is general information only, does not constitute legal advice, and no lawyer-client relationship is created by it. The cases discussed remain subject to appeal and further proceedings, and related litigation is ongoing in multiple jurisdictions; the law in this area is unsettled and evolving rapidly. Specific advice from qualified counsel should be taken on any particular AI, data-sourcing or intellectual-property question before any decision is made.

© 2026 Muzy & Meraris LLP. All rights reserved.

Navigate

Contact

Connect

contact@muzylaw.com

+924232216907