A generative artificial intelligence model trained on public-domain artwork from 744 different artists produced a single portrait. Researchers then simulated removing each of those 744 artists from the training data, one at a time, and generated the image again in each scenario. Removing the data changed almost nothing.
According to a study published in Nature Communications, nearly all of the versions of the generated image looked almost identical to the original.
Download the Straight Arrow app today to get the stories that matter free from manipulation, bias or agenda.™
Point phone camera here
Researchers at the Massachusetts Institute of Technology’s Computer Science and Artificial Intelligence Laboratory call this effect “attribution decay” — what lead author Zheng Dai described as “the phenomenon that larger data sets induce lower attributability.” The finding is the first to prove the effect directly rather than estimate it, a distinction that matters because every earlier attempt at answering this question relied on approximation.

How did the study work?
Previous attempts to trace an AI image back to its training data relied on approximations rather than proof. Testing the question directly would mean actually removing a piece of training data and checking what changes. But that would normally require retraining an entire model from scratch — for every single artist or image a researcher wanted to check.
That’s far too costly to do at any meaningful scale. So earlier researchers used mathematical shortcuts instead. They estimated an image’s likely influence rather than truly isolating it.
Dai and his co-author, MIT professor David Gifford, found a way around that limitation.
Normally, a model learns from all its training data at once, making it impossible to cleanly separate any single influence later. Instead, Dia and Gifford split the model into many smaller components. Each one trained independently on only a slice of data. Generating an image means combining all those components’ outputs together.
To test what would happen without a specific artist or image, the researchers simply pulled out the handful of training components, then combined the rest. This à la carte method didn’t require retraining. They call this structure a “diffusion ensemble.”
Dai has a theory for why removing one piece so rarely changes the outcome. The same visual patterns tend to show up repeatedly across a large dataset. So no single image or artist is the only place a given detail could have come from.
In the simplest version of that idea, he said, two nearly identical images might sit in the same dataset. Remove one, and the pattern survives in the other.
“If you have two of the same images in the training set, you could remove one of them, and it probably wouldn’t do much,” he told Straight Arrow.
Across a dataset of hundreds of thousands of images, he suspects that kind of overlap happens constantly, just less obviously.
A question the study didn’t answer
A natural follow-up question is what happens when someone doesn’t let a model generate freely but specifically asks it to create an image in the style of a specific artist. The study didn’t test that scenario.
“This is different from what we measured,” he told Straight Arrow. “We didn’t measure with prompts that specifically say, ‘Draw this in the style of Monet.’”
Instead, the study tested unprompted generation, or letting the model produce an image on its own. Then researchers checked whether removing the artist behind the closest-looking training image changed the result.
Dai flagged one important nuance: an image can visually resemble a specific artist’s work without technically being “attributable” to that artist in the sense the study measured.
“It is possible that it generates something that visually looks like something by Monet, but actually is not attributable to Monet,” he said.
James Grimmelmann, a professor at Cornell Law School and Cornell Tech who studies AI and copyright law, said that gap matters because “style” claims raise a different legal question entirely.
No artist owns generic techniques, he said. For example, nobody could claim exclusive rights over two eyes stacked vertically on one side of a head, a nod to Pablo Picasso’s cubist portraits. The real dispute, he said, is “not really about whether you sourced from them at all. It’s how much did you take.”
What’s at stake?
The stakes go beyond how AI models work internally. A derivative work is one that builds on, adapts or transforms something that already exists. Legally, an unauthorized one can amount to copyright infringement.
But to prove that in court, someone typically has to show the alleged copier had access to the original work first. Attribution, if it had worked reliably, would have given courts a direct way to check that. Grimmelmann described what that would’ve looked like: “These similarities in this output are genuinely due to the influence of this work in the training data.” That’s exactly the tool this study suggests doesn’t hold up once a model is trained at scale.
The study’s authors go further than Grimmelmann does. In their discussion, they write that unattributability provides a “refutation of access.” That’s the exact requirement described above.
In practical terms, once a model is trained on enough data, proving that a specific copyrighted work influenced a specific output may become almost impossible. That’s even true if the work was actually used in training.
Grimmelmann said the finding doesn’t clearly favor either side of these disputes, saying, “It complicates things for both sides.” Had attribution worked, artists could’ve pointed to real copying when it happened. AI companies could’ve pointed to clean synthesis when it didn’t. Without that tool, cases don’t resolve as neatly in either direction.
Research like this isn’t staying in academic journals, either. Grimmelmann noted that German courts have already drawn on similar studies of AI memorization. The finding lands in the middle of active litigation, including the $1.5 billion settlement in Bartz v. Anthropic and the ongoing Andersen v. Stability AI case.
Grimmelmann described copyright cases as a three-stage pipeline.
First, a court has to establish that copying happened at all. Then it checks whether there was too much of it. Finally, it decides whether the copying was justified under fair use. Weakening the first stage doesn’t end the first, he said; it just “leaves substantial similarity and fair use with more work to do.”
What the study doesn’t say
Grimmelmann cautioned against reading too deeply into the finding.
“If I was reading a headline about this and didn’t look deeper, I might think it means attribution can never work in any setting,” he said, which isn’t what the study is claiming.
He also worried it could reinforce “a perception of AI systems as uncontrollable black boxes that nobody understands.” Even without precise attribution, he said, companies aren’t powerless to control what their models produce regarding copyright. Grimmelmann pointed to techniques like reinforcement learning that train models to refuse requests to closely duplicate existing work.
The study itself includes two caveats worth noting. First, exact copying hasn’t disappeared entirely. Even large, commercially deployed models still occasionally generate near-identical reproductions of a training image. However, the researchers note that prior work puts this at about one in a million outputs. Second, the same effect that makes artists’ work harder to trace also applies to photos of real people. This means attribution decay can work as an incidental privacy protection, not just a compilation for copyright claims.
Asked where this leaves the broader debate over whether training AI on copyrighted work amounts to infringement, Grimmelmann didn’t predict a resolution.
“It was messy before,” he said. “It’s messy afterwards.”
The study doesn’t settle whether AI companies broke the law by training on artists’ work; it just makes that question harder to answer in court.
Round out your reading
- All your questions about napping, answered.
- Trump’s $5,000 promise echoes past payouts that never materialized.
- Inside the effort to make data centers pay their share of electricity costs.
- Why did the Feds seize the ’largest Martian meteorite on Earth’?
- Photos and video show exactly where and when Trump visited Ground Zero after 9/11.