Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.
If you request an AI agent to hack an account, it will likely decline, but researchers at EPFL have demonstrated a simpler method that relies on patience rather than technical expertise. Their recent study indicates that dividing a harmful objective into smaller, seemingly harmless requests can deceive AI agents into executing tasks they would typically refuse (according to TechXplore).
This concept mirrors the recent 'Bioshocking' exploit where AI browsers were tricked into perceiving credential theft as part of an innocuous game.
How the researchers revealed this vulnerability
The team developed an automated testing tool named STING, which stands for Sequential Testing of Illicit N-step Goal execution, intended to simulate the way a genuine attacker would operate. Rather than directly stating a harmful goal, STING strategically breaks it down into a series of smaller, seemingly benign steps that lead up to the objective over several conversational exchanges.
Researchers assessed this method across 176 harmful scenarios using prominent AI models like ChatGPT, Gemini, and Claude. Each AI agent was evaluated as a tool-using agent, capable of web browsing, sending emails, and accomplishing multistep tasks.
Gradual, step-by-step manipulation proved to be significantly more successful than straightforward, single-prompt attempts. In certain instances, models were twice as likely to carry out a harmful task when the request was divided into smaller components. This finding aligns with other research indicating that even typical users can circumvent AI safety measures with carefully crafted prompts.
Why is this significant?
The concerns highlighted by the researchers are not merely theoretical risks. In June, Meta acknowledged that attackers utilized simple social engineering rather than malware or hacking tools to deceive its AI support assistant into granting unauthorized access to Instagram accounts.
The researchers initially expected attacks to be more effective in languages with limited training data. However, they discovered that completion rates were relatively consistent across all seven tested languages. They found one exception: switching languages during a multi-step attack significantly increased success rates.
Lead researcher Ayush Kumar Tarun believes that safety testing should be incorporated much earlier in the design of an agent. Simply adding it on after an issue arises is inadequate, especially as these systems continue to acquire more practical capabilities.
---
Meta has finally introduced its AI assistant to the Mac with a dedicated desktop application, providing users with yet another AI chatbot to keep open alongside other tabs. The app is capable of analyzing a shared window, answering questions related to what's displayed on the screen, generating content, and processing voice input across macOS.
With the Mac now equipped to have its AI assistant visually recognize content displayed on-screen.
---
Avec, the AI-driven email application utilizing swipe gestures and voice dictation, has now implemented a feature that ensures you never overlook a deadline again. We have all experienced the situation where a deadline gets lost in a lengthy email thread, life becomes hectic, and before you know it, you're apologizing for forgetting a follow-up. Avec has rolled out a solution for this issue, which may be its most intelligent enhancement yet.
---
I have long been a Mac user, and over the years, I've become proficient at diagnosing issues when something seems amiss. Typically, I begin by exploring the Settings app, making minor adjustments, and usually identify the problem. Storage issues are a prime example; I recently observed my Mac slowing down, accessed the Storage section, and immediately saw a clear breakdown of what was consuming space. After a quick cleanup, I reclaimed valuable storage, and everything returned to normal. I've also applied this method to investigate battery drainage or unexpected CPU usage. macOS offers various tools to examine when something isn't functioning correctly.
However, one aspect it doesn’t highlight clearly is where your internet data is actually being utilized. I began pondering this issue when my internet had been acting unusually over the past few days. Everything seemed slower than usual, yet I couldn't pinpoint the cause. My connection appeared fine, and I wasn’t downloading large files, so what was consuming all that bandwidth? It turned out several applications on my Mac were using the internet in the background, which I was unaware of. I likely wouldn't have figured this out if I hadn't stumbled upon a Mac app that revealed exactly what was happening behind the scenes. That’s when things started to get intriguing.
Other articles
Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.
This recent research shows how careful, incremental adjustments can deceive AI agents into overlooking their own safety protocols.
