AI could be acquiring knowledge from countless images without replicating any of them.
MIT researchers have discovered what they term “attribution decay” in large generative models.
Researchers created an image based on artwork from 744 artists and then compared the outputs obtained when each artist was individually excluded from the training data.
Here’s a challenging question for the generative AI age: if an AI produces an image, can anyone truly identify the specific images that influenced it? A recent study from MIT’s Computer Science and Artificial Intelligence Laboratory indicates that for sufficiently large models, the answer may often be no. Researchers Dai and Gifford describe this phenomenon as “attribution decay” in their open-access paper, published today in Nature Communications. It highlights how the impact of individual training data becomes harder to trace as the dataset enlarges.
The larger the dataset, the more ambiguous the attribution becomes.
The researchers aimed to investigate something more specific than merely determining whether an AI model had been trained on a certain image. Their method was fundamentally counterfactual: what occurs if a specific piece of training data is left out?
Eliminating one artist’s work significantly altered outcomes in a smaller dataset, but had little effect with a 50,000-image dataset, demonstrating how attribution diminishes as datasets expand.
They discovered that, as models and datasets increase in size, the removal of an individual image often results in little or no noticeable change to the generated output. The same applies to removing the entire works of an artist or images of a specific individual. In essence, the model may have absorbed broad visual patterns from a vast array of data without any single image being clearly accountable for a specific output. MIT researchers Zheng Dai and Professor David Gifford assert that if removing a specific piece of data does not alter the output, it becomes challenging to meaningfully attribute that output to the excluded example.
This does not resolve the AI copyright issue.
It should be noted that this study does not imply that training data is insignificant, nor does it settle the larger debate about whether AI companies may utilize copyrighted materials without consent. Instead, its focus is more specific: whether a particular training image can be demonstrated to have directly influenced a specific AI-generated result. A model may not directly replicate any individual artist’s work while still drawing on millions of images to learn aspects like composition, lighting, textures, and artistic styles.
Midjourney is a versatile, high-quality AI art generator, as highlighted in the community showcase.
A more intriguing takeaway is that this connection becomes increasingly hard to track as datasets expand. An AI-produced image may rely on patterns acquired from an extensive collection of material without having a clear, identifiable source image behind it. The model may have learned from a wide range, while no single image unmistakably leaves a mark on the final result.
Varun is a seasoned technology journalist and editor with over eight years of experience in consumer tech media. His work encompasses…
Meta's AI assistant has finally arrived on Mac, though it has some ground to cover.
The new desktop assistant can analyze visible content on the screen and accept voice commands throughout macOS.
Meta has successfully launched its Meta AI assistant on the Mac with a dedicated desktop application, providing users with yet another AI chatbot to keep alongside their many tabs. The app can analyze a shared window, respond to inquiries about the on-screen content, generate text, and accept voice dictation across macOS.
Meta AI can actually observe what's on the screen.
Researchers are developing straightforward 3D-printed reflective surfaces to reroute mmWave signals in challenging indoor spaces.
5G's millimeter-wave technology offers incredibly swift wireless speeds but has a notable drawback: those signals struggle to navigate around obstacles. Walls, furniture, people, and even minor movements can significantly weaken mmWave connections. Now, researchers at the University of California San Diego are suggesting a surprisingly simple solution: inexpensive reflective tiles that effectively direct the signal where it's needed.
The challenge of exceptionally fast 5G.
Regardless of personal views, ChatGPT is already ubiquitous. People utilize it to investigate complex subjects, refine their writing, learn new skills, organize their lives, and, in my case, ask an embarrassing number of questions I could have figured out myself. Teenagers are no exception. In fact, OpenAI reports that nearly 90% of teens who use ChatGPT do so to learn, gather information, develop skills, or accomplish tasks in a typical week. Hence, rather than pretending students won’t use AI, OpenAI is embracing a different strategy: providing them with a version specifically designed for them. The company has announced ChatGPT for Teens, a new experience tailored for users aged 13 to 17 that focuses on learning, promoting healthier usage habits, and implementing significantly more robust safety measures.
ChatGPT seeks to educate, not merely provide answers.
Other articles
AI could be acquiring knowledge from countless images without replicating any of them.
A study by MIT reveals that images generated by AI frequently cannot be linked to particular training images, since larger datasets diminish the impact of specific examples.
