AI could be acquiring knowledge from billions of images without replicating any individual one.
MIT researchers have discovered a phenomenon they refer to as “attribution decay” within large generative models.
The study involved generating an image using artwork by 744 artists, which was then compared to outputs that resulted from the exclusion of individual artists from the training dataset.
A pressing question in the realm of generative AI arises: if an AI generates an image, can anyone accurately identify the specific images that influenced it? Recent research from MIT’s Computer Science and Artificial Intelligence Laboratory indicates that, for sufficiently large models, the answer may frequently be no. Researchers Zheng Dai and David Gifford define this phenomenon as “attribution decay” in their open-access article published in Nature Communications. It describes how the impact of any single training data point becomes increasingly elusive as the size of the dataset increases.
As datasets expand, the connection between specific images and their influence becomes less clear.
The researchers aimed to investigate more precisely than simply determining whether an AI model was trained on a specific image. Their method was essentially counterfactual: what occurs when a particular piece of training data is omitted?
Removing one artist’s work led to significant changes in outputs from a smaller dataset, but had minimal effect when applied to a dataset containing 50,000 images, demonstrating how attribution diminishes as datasets increase in size.
Their findings suggest that as models and datasets grow, the removal of an individual image may result in negligible or no measurable difference in the generated output. This same principle applies when an entire artist’s works or images of a specific individual are omitted. Thus, the model may have assimilated broad visual patterns from an extensive data pool without any single image being distinctly accountable for a particular output. According to MIT researchers Dai and Gifford, if omitting a specific data piece does not alter the output, it becomes challenging to meaningfully attribute that output to the excluded item.
However, this does not resolve the ongoing AI copyright discussion.
It's important to clarify that the study does not imply that training data is insignificant, nor does it settle the larger debate over whether AI companies can utilize copyrighted material without authorization. The focus is more limited: it examines whether a particular training image can be demonstrated to have directly influenced a specific AI-generated output. A model may not replicate any single artist’s work yet still leverage millions of images to acquire knowledge about composition, lighting, textures, and artistic styles.
Midjourney is a flexible, high-quality AI art generator, as evidenced in the community showcase.
The more intriguing insight is that this connection becomes harder to trace as datasets expand. An AI-generated image may reflect patterns learned from a vast array of materials without a clear, identifiable source image associated with it. The model could have absorbed influences from many sources, yet no single image may leave an obvious mark on the final outcome.
Other articles
AI could be acquiring knowledge from billions of images without replicating any individual one.
A study conducted by MIT reveals that AI-generated images frequently cannot be linked back to specific training images, since larger datasets diminish the impact of individual examples.
