If your art trained the AI, MIT says you may never be able to prove it


Full story

A generative artificial intelligence model trained on public-domain artwork from 744 different artists produced a single portrait. Researchers then simulated removing each of those 744 artists from the training data, one at a time, and generated the image again in each scenario. Removing the data changed almost nothing.

According to a study published in Nature Communications, nearly all of the versions of the generated image looked almost identical to the original.

QR code for SAN app download

Download the Straight Arrow app today to get the stories that matter free from manipulation, bias or agenda.™

Point phone camera here

Researchers at the Massachusetts Institute of Technology’s Computer Science and Artificial Intelligence Laboratory call this effect “attribution decay” — what lead author Zheng Dai described as “the phenomenon that larger data sets induce lower attributability.” The finding is the first to prove the effect directly rather than estimate it, a distinction that matters because every earlier attempt at answering this question relied on approximation. 

The top left of the image is an image generated by a model trained on public domain artwork created by 744 artists. The others are images that would’ve been generated had any one of the 744 artists been omitted from the training set. (Credit: Dai & Gifford, Nature Communications, 2026)
The top left of the image is an image generated by a model trained on public domain artwork created by 744 artists. The others are images that would’ve been generated had any one of the 744 artists been omitted from the training set. (Credit: Dai & Gifford, Nature Communications, 2026)

How did the study work?

Previous attempts to trace an AI image back to its training data relied on approximations rather than proof. Testing the question directly would mean actually removing a piece of training data and checking what changes. But that would normally require retraining an entire model from scratch — for every single artist or image a researcher wanted to check. 

That’s far too costly to do at any meaningful scale. So earlier researchers used mathematical shortcuts instead. They estimated an image’s likely influence rather than truly isolating it. 

Dai and his co-author, MIT professor David Gifford, found a way around that limitation. 

Normally, a model learns from all its training data at once, making it impossible to cleanly separate any single influence later. Instead, Dia and Gifford split the model into many smaller components. Each one trained independently on only a slice of data. Generating an image means combining all those components’ outputs together.

To test what would happen without a specific artist or image, the researchers simply pulled out the handful of training components, then combined the rest. This à la carte method didn’t require retraining. They call this structure a “diffusion ensemble.”

Dai has a theory for why removing one piece so rarely changes the outcome. The same visual patterns tend to show up repeatedly across a large dataset. So no single image or artist is the only place a given detail could have come from. 

In the simplest version of that idea, he said, two nearly identical images might sit in the same dataset. Remove one, and the pattern survives in the other.

“If you have two of the same images in the training set, you could remove one of them, and it probably wouldn’t do much,” he told Straight Arrow.

Across a dataset of hundreds of thousands of images, he suspects that kind of overlap happens constantly, just less obviously. 

A question the study didn’t answer

A natural follow-up question is what happens when someone doesn’t let a model generate freely but specifically asks it to create an image in the style of a specific artist. The study didn’t test that scenario.

“This is different from what we measured,” he told Straight Arrow. “We didn’t measure with prompts that specifically say, ‘Draw this in the style of Monet.’”

Instead, the study tested unprompted generation, or letting the model produce an image on its own. Then researchers checked whether removing the artist behind the closest-looking training image changed the result. 

Dai flagged one important nuance: an image can visually resemble a specific artist’s work without technically being “attributable” to that artist in the sense the study measured. 

“It is possible that it generates something that visually looks like something by Monet, but actually is not attributable to Monet,” he said.

James Grimmelmann, a professor at Cornell Law School and Cornell Tech who studies AI and copyright law, said that gap matters because “style” claims raise a different legal question entirely. 

No artist owns generic techniques, he said. For example, nobody could claim exclusive rights over two eyes stacked vertically on one side of a head, a nod to Pablo Picasso’s cubist portraits. The real dispute, he said, is “not really about whether you sourced from them at all. It’s how much did you take.”

What’s at stake? 

The stakes go beyond how AI models work internally. A derivative work is one that builds on, adapts or transforms something that already exists. Legally, an unauthorized one can amount to copyright infringement. 

But to prove that in court, someone typically has to show the alleged copier had access to the original work first. Attribution, if it had worked reliably, would have given courts a direct way to check that. Grimmelmann described what that would’ve looked like: “These similarities in this output are genuinely due to the influence of this work in the training data.” That’s exactly the tool this study suggests doesn’t hold up once a model is trained at scale. 

The study’s authors go further than Grimmelmann does. In their discussion, they write that unattributability provides a “refutation of access.” That’s the exact requirement described above. 

In practical terms, once a model is trained on enough data, proving that a specific copyrighted work influenced a specific output may become almost impossible. That’s even true if the work was actually used in training

Grimmelmann said the finding doesn’t clearly favor either side of these disputes, saying, “It complicates things for both sides.” Had attribution worked, artists could’ve pointed to real copying when it happened. AI companies could’ve pointed to clean synthesis when it didn’t. Without that tool, cases don’t resolve as neatly in either direction. 

Research like this isn’t staying in academic journals, either. Grimmelmann noted that German courts have already drawn on similar studies of AI memorization. The finding lands in the middle of active litigation, including the $1.5 billion settlement in Bartz v. Anthropic and the ongoing Andersen v. Stability AI case. 

Grimmelmann described copyright cases as a three-stage pipeline. 

First, a court has to establish that copying happened at all. Then it checks whether there was too much of it. Finally, it decides whether the copying was justified under fair use. Weakening the first stage doesn’t end the first, he said; it just “leaves substantial similarity and fair use with more work to do.” 

What the study doesn’t say

Grimmelmann cautioned against reading too deeply into the finding.

“If I was reading a headline about this and didn’t look deeper, I might think it means attribution can never work in any setting,” he said, which isn’t what the study is claiming. 

He also worried it could reinforce “a perception of AI systems as uncontrollable black boxes that nobody understands.” Even without precise attribution, he said, companies aren’t powerless to control what their models produce regarding copyright. Grimmelmann pointed to techniques like reinforcement learning that train models to refuse requests to closely duplicate existing work. 

The study itself includes two caveats worth noting. First, exact copying hasn’t disappeared entirely. Even large, commercially deployed models still occasionally generate near-identical reproductions of a training image. However, the researchers note that prior work puts this at about one in a million outputs. Second, the same effect that makes artists’ work harder to trace also applies to photos of real people. This means attribution decay can work as an incidental privacy protection, not just a compilation for copyright claims. 

Asked where this leaves the broader debate over whether training AI on copyrighted work amounts to infringement, Grimmelmann didn’t predict a resolution.

“It was messy before,” he said. “It’s messy afterwards.” 

The study doesn’t settle whether AI companies broke the law by training on artists’ work; it just makes that question harder to answer in court.

Round out your reading

Tags: , , , ,

Straight Arrow
Fear No Fact.

Don't just take our word for it.


Center-rated reporting

According to media bias experts at AllSides

AllSides Center-rated reporting May 2026

Transparent and credible

Awarded a perfect reliability rating from NewsGuard

100/100

Welcome back to trustworthy journalism.

Find out more

Why this story matters

A peer-reviewed MIT study finds that once an AI image model is trained on enough data, tracing a specific output back to any individual artist's work becomes nearly impossible — a finding now entering active copyright litigation.

Copyright claims get harder to prove

According to the study's authors, unattributability undermines the legal requirement to show a specific copyrighted work influenced a specific AI output, even if that work was used in training.

Active lawsuits are affected

The finding lands amid ongoing cases, including Andersen v. Stability AI and a $1.5 billion settlement in Bartz v. Anthropic, where establishing that copying occurred is a required first step.

Style prompts remain untested

The study did not measure what happens when users specifically prompt a model to generate images in a named artist's style, leaving that legally distinct question unanswered.

Straight Arrow
Fear No Fact.

Don't just take our word for it.


Center-rated reporting

According to media bias experts at AllSides

AllSides Center-rated reporting May 2026

Transparent and credible

Awarded a perfect reliability rating from NewsGuard

100/100

Welcome back to trustworthy journalism.

Find out more