philip lelyveld The world of entertainment technology

24Aug/26Off

When AI art has no author: MIT study finds generated images often can’t be traced to any training data

PhilNote: we're back to "correlation versus causation"

... at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.

And if removing something changes nothing, the researchers argue, it can't be said to be responsible for anything. "If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output," says Zheng Dai SM ‘21, PhD ‘24, former MIT CSAIL researcher and lead author on the work. ...

"All previous methods were approximate," says MIT professor and MIT CSAIL principal investigator David Gifford. "They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute. You're actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don't change." ...

Exploring a counterfactual universe

With ablation working, the researchers could finally ask their question at scale. Take one generated image, then imagine every alternate version of it, each produced by removing a different piece of the training data. The team calls this the image's counterfactual universe. The distance between the original and its most different alternate, the counterfactual radius, captures the most any single piece of training data could have mattered. ...

The privacy paradox

The implications run in a direction that surprised the researchers themselves.

Gifford sees the finding as bearing directly on the legal question of whether model outputs are derivative works. “One way to think about this is that these models are creative. They are not simply copying what they are fed, but creating brand new outputs. If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn't attributable to anything on the internet.”  ...

...this paper provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying.” ...

See the full story here: https://www.csail.mit.edu/news/when-ai-art-has-no-author-mit-study-finds-generated-images-often-cant-be-traced-any-training

Comments (0) Trackbacks (0)

Sorry, the comment form is closed at this time.

Trackbacks are disabled.