This experiment demonstrates how simple it is to compromise an open-weight AI model for less than $100.
This research introduces new skepticism regarding the trustworthiness of open-weight AI models.
Recently, open-weight AI models have gained significant attention. For instance, Moonshot’s extensive Kimi K3 model has performed closely to Claude Fable 5 and GPT 5.6 Sol in various benchmarks, while remaining entirely open-weight and accessible for anyone to download.
However, Katie Paxton-Fear, a cybersecurity lecturer at Manchester Metropolitan University and a security advocate at Semgrep, has managed to poison an open-weight model, illustrating how easily that openness can be exploited (as reported by The Register).
How did the researcher poison the AI model so rapidly?
Paxton-Fear began with a small experiment to check if fine-tuning could covertly alter a model's adherence to JavaScript coding conventions, even when explicitly instructed otherwise. After her initial experiment succeeded with minimal resistance, she chose to go further and create a backdoor in the model.
I started out by trying to figure out if I could use fine-tuning to get a model to switch from camelCase for JavaScript to snake_case, and it was surprisingly simple, even after we instructed the AI to use camelCase. After that worked, I created a proper backdoor. pic.twitter.com/35alEwypn8— Katie Paxton-Fear (@InsiderPhD) July 14, 2026
It took just ten poisoned training examples for the model to begin reliably generating code that was vulnerable to remote code execution, a vulnerability that allows attackers to execute their commands on another person's machine.
The entire process cost less than $100 and took approximately an hour. Notably, larger AI models were found to be even easier to compromise than smaller ones, which aligns with a pattern identified in a University of Washington study where more powerful AI browsers posed the highest security threats among those assessed.
Why is this concerning for users of open-weight models?
The primary worry is not just the potential for a model to be poisoned, but rather that there are few dependable methods for detecting whether it has been altered. Traditional software can be reverse-engineered to fully understand its behavior, but AI models do not provide the same level of transparency, even if they are open-weight.
So can we trust open-weight models that are fine-tuned online and marketed as the solution to our AI token spending issues? It seems we may need more than just benchmarks and warnings to avoid writing insecure code— Katie Paxton-Fear (@InsiderPhD) July 14, 2026
A compromised model doesn't need to noticeably fail to inflict harm; it only needs to subtly influence decisions in ways that go unnoticed. Commercial closed models like Claude or ChatGPT are also not exempt, as they require a significant amount of trust while providing minimal insight into their workings. This research serves as a clear reminder that placing unreserved trust in any AI model, regardless of being open-weight or not, carries real risks.
Manisha Priyadarshini is a tech and entertainment writer with over nine years of editorial experience.
Samsung's latest OLED laptop panels are now brighter and more durable.
Featuring peak brightness of 1,600 nits, deeper blacks, and a doubled lifespan, all within one panel. A month after showcasing its brightest smartphone displays, Samsung Display is now enhancing laptop screens. The company announced today that it has started supplying a new tandem OLED panel for laptops that achieves the new DisplayHDR True Black 1400 certification by VESA, the highest brightness standard that OLED panels have had to meet thus far. If you've ever found your laptop display lacking in brightness outdoors, this update is specifically aimed at you.
What improvements does True Black 1400 offer?
Apple's Mac Pro was almost equipped with an M3 Extreme chip that was twice as powerful as the M3 Ultra.
High production costs likely derailed Apple’s plans for the M3 Extreme chip.
Earlier this year, Apple discontinued the Mac Pro, ending a 20-year legacy for a computer that once epitomized the best of the company's desktop offerings. However, reports indicate that Apple had even grander ambitions for the device before ultimately replacing it with the Mac Studio. According to Bloomberg’s Mark Gurman, Apple had developed an M3 Extreme chip that could have provided twice the CPU and GPU cores compared to the M3 Ultra. This processor was intended to be positioned above the Ultra tier and could have finally given the Mac Pro the performance edge it sorely lacked. Ultimately, Apple scrapped the chip due to concerns regarding production costs and a limited market demand for such a high-priced machine.
Researchers have found that hidden prompts can secretly alter an AI’s memory, posing a serious issue.
Recent studies reveal an AI attack capable of rewriting an assistant's long-term memory.
Large language models are increasingly enhancing their ability to remember user preferences, such as writing styles, recurring tasks, shopping habits, and project deadlines, making interactions feel more personalized and effective. However, new findings suggest that this very feature
Other articles
This experiment demonstrates how simple it is to compromise an open-weight AI model for less than $100.
A cybersecurity researcher successfully poisoned an open-weight AI model for less than $100 in approximately one hour, highlighting how effortlessly these growingly popular systems can be covertly compromised without being noticed.
