Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.

Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.

      If you request an AI agent to hack an account, it will likely decline, but researchers at EPFL have demonstrated a simpler method that relies on patience rather than technical expertise. Their recent study indicates that dividing a harmful objective into smaller, seemingly harmless requests can deceive AI agents into executing tasks they would typically refuse (according to TechXplore).

      This concept mirrors the recent 'Bioshocking' exploit where AI browsers were tricked into perceiving credential theft as part of an innocuous game.

      How the researchers revealed this vulnerability

      The team developed an automated testing tool named STING, which stands for Sequential Testing of Illicit N-step Goal execution, intended to simulate the way a genuine attacker would operate. Rather than directly stating a harmful goal, STING strategically breaks it down into a series of smaller, seemingly benign steps that lead up to the objective over several conversational exchanges.

      Researchers assessed this method across 176 harmful scenarios using prominent AI models like ChatGPT, Gemini, and Claude. Each AI agent was evaluated as a tool-using agent, capable of web browsing, sending emails, and accomplishing multistep tasks.

      Gradual, step-by-step manipulation proved to be significantly more successful than straightforward, single-prompt attempts. In certain instances, models were twice as likely to carry out a harmful task when the request was divided into smaller components. This finding aligns with other research indicating that even typical users can circumvent AI safety measures with carefully crafted prompts.

      Why is this significant?

      The concerns highlighted by the researchers are not merely theoretical risks. In June, Meta acknowledged that attackers utilized simple social engineering rather than malware or hacking tools to deceive its AI support assistant into granting unauthorized access to Instagram accounts.

      The researchers initially expected attacks to be more effective in languages with limited training data. However, they discovered that completion rates were relatively consistent across all seven tested languages. They found one exception: switching languages during a multi-step attack significantly increased success rates.

      Lead researcher Ayush Kumar Tarun believes that safety testing should be incorporated much earlier in the design of an agent. Simply adding it on after an issue arises is inadequate, especially as these systems continue to acquire more practical capabilities.

      ---

      Meta has finally introduced its AI assistant to the Mac with a dedicated desktop application, providing users with yet another AI chatbot to keep open alongside other tabs. The app is capable of analyzing a shared window, answering questions related to what's displayed on the screen, generating content, and processing voice input across macOS.

      With the Mac now equipped to have its AI assistant visually recognize content displayed on-screen.

      ---

      Avec, the AI-driven email application utilizing swipe gestures and voice dictation, has now implemented a feature that ensures you never overlook a deadline again. We have all experienced the situation where a deadline gets lost in a lengthy email thread, life becomes hectic, and before you know it, you're apologizing for forgetting a follow-up. Avec has rolled out a solution for this issue, which may be its most intelligent enhancement yet.

      ---

      I have long been a Mac user, and over the years, I've become proficient at diagnosing issues when something seems amiss. Typically, I begin by exploring the Settings app, making minor adjustments, and usually identify the problem. Storage issues are a prime example; I recently observed my Mac slowing down, accessed the Storage section, and immediately saw a clear breakdown of what was consuming space. After a quick cleanup, I reclaimed valuable storage, and everything returned to normal. I've also applied this method to investigate battery drainage or unexpected CPU usage. macOS offers various tools to examine when something isn't functioning correctly.

      However, one aspect it doesn’t highlight clearly is where your internet data is actually being utilized. I began pondering this issue when my internet had been acting unusually over the past few days. Everything seemed slower than usual, yet I couldn't pinpoint the cause. My connection appeared fine, and I wasn’t downloading large files, so what was consuming all that bandwidth? It turned out several applications on my Mac were using the internet in the background, which I was unaware of. I likely wouldn't have figured this out if I hadn't stumbled upon a Mac app that revealed exactly what was happening behind the scenes. That’s when things started to get intriguing.

Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures. Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.

Other articles

Researchers have highlighted a concerningly straightforward method that allows AI bots to bypass safety protocols and act unpredictably. Researchers have highlighted a concerningly straightforward method that allows AI bots to bypass safety protocols and act unpredictably. This recent study demonstrates how carefully staged manipulation can deceive AI agents into bypassing their own safety protocols. Meta introduced a Mac application designed to implement its AI technology for businesses. Meta introduced a Mac application designed to implement its AI technology for businesses. Meta has launched a Mac application for Meta AI, aimed at creators and small businesses. The app integrates with Instagram, Facebook, and advertising accounts to provide suggestions on what to post. Samsung has raised foundry prices by as much as 15%, with China facing the highest increases. Samsung has raised foundry prices by as much as 15%, with China facing the highest increases. According to Reuters, Samsung has increased its advanced chipmaking prices by as much as 15%, driven by AI demand exceeding TSMC’s production capacity. Chinese clients are now facing the highest prices. Just five days after acquiring Cursor for $60 billion, SpaceX attempted to purchase Cognition. Just five days after acquiring Cursor for $60 billion, SpaceX attempted to purchase Cognition. According to Bloomberg, SpaceX attempted to acquire the AI coding startup Cognition. However, the CEO stated that the company is not for sale and that no discussions took place. Smart glasses priced at $100 can now observe, hear, and translate the surroundings. Smart glasses priced at $100 can now observe, hear, and translate the surroundings. A pair of AI smart glasses priced at $99.99 features a 12MP Sony camera, visual recognition, live translation, open-ear audio, and voice assistance, all within unremarkable-looking frames. According to reports, PlayStation is revamping Horizon Hunters Gathering following unsatisfactory playtest results. According to reports, PlayStation is revamping Horizon Hunters Gathering following unsatisfactory playtest results. PlayStation is said to be revamping Horizon Hunters Gathering, removing its live-service model due to unsatisfactory playtests, and employees are under pressure to meet a deadline in December.

Researchers reveal an alarmingly straightforward method to make AI bots behave unpredictably and bypass safety measures.

This recent research shows how careful, incremental adjustments can deceive AI agents into overlooking their own safety protocols.