Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.

Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.

      Large language models are improving their ability to remember user-specific details. Whether it's an individual's writing style, regular tasks, shopping preferences, or project deadlines, AI assistants are increasingly capable of storing long-term memories, making interactions feel more personal and beneficial. However, recent research indicates that this very capability might turn into one of AI’s greatest security weaknesses.

      Researchers from New Mexico State University have introduced a new tactic known as GhostWriter, which can secretly insert false memories into AI agents. Instead of outright information theft, this attack manipulates an AI's memory, potentially leading to hazardous decisions long after the initial attack.

      This marks a subtle yet crucial change in how AI systems can be compromised. Rather than targeting the AI model directly, attackers focus on its memory.

      The attack does not involve hacking the AI; rather, it alters what the AI remembers.

      Traditional chatbots typically have minimal or no memory across interactions. On the other hand, contemporary AI agents increasingly depend on persistent memory systems that gather data about users, ongoing projects, and past interactions, enabling them to offer more contextual and personalized replies over time.

      The researchers claim that these memory systems introduce a completely new attack surface. GhostWriter operates by subtly injecting harmful information into an AI agent’s long-term memory via concealed prompts or unreliable external content. This false information remains inactive until the AI retrieves it while responding to an otherwise valid request.

      For instance, if you ask your AI assistant to summarize emails from your bank, a compromised memory could result in the AI secretly forwarding those emails to an attacker instead. It might also retain incorrect contact details, fake deadlines, wrong preferences, or fabricated information, all due to someone successfully altering the assistant’s perceived truths.

      Unlike traditional prompt injection attacks, which typically impact a single conversation, GhostWriter is designed for persistence. Once harmful information enters the memory, it can continue to affect the AI’s behavior across multiple future sessions until it is identified and eliminated.

      The researchers outline the attack as a two-phase process. The first phase involves memory injection, where harmful content is silently stored in the AI's memory. The second phase is attack activation, which occurs when the AI unknowingly retrieves that tainted memory while responding to a legitimate user query.

      As AI memory becomes increasingly useful, it also becomes an attractive target for attacks.

      The timing of this research is notable. Nearly all major AI companies are striving to create assistants that retain user information over weeks, months, or even years. Memory has rapidly turned into one of the industry's key differentiators, as it enables AI to feel less like a basic chatbot and more like a personal assistant.

      However, this also means that memory requires protection comparable to the model itself. In their experiments, the researchers found that GhostWriter achieved a memory injection success rate of approximately 98%, with malicious memories being activated about 60% of the time against cutting-edge AI agents. These figures imply that current memory architectures may not yet be able to differentiate between trustworthy information and manipulated data.

      The research team is not only addressing the problem but has also proposed a protective framework called Agentic Memory Sentry (AM-Sentry). This framework integrates memory screening with stricter memory management policies. The researchers report that this approach significantly decreased GhostWriter's success rate while maintaining the AI's utility.

      As AI agents progress into digital assistants capable of managing emails, scheduling meetings, writing code, and making decisions for users, ensuring the security of their memories may become as crucial as protecting the information they produce. The next challenge in AI security might not just be safeguarding models from harmful prompts, but also defending their memories against being rewritten entirely.

Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue. Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.

Other articles

AliExpress faces a significant fine: the EU has imposed a record €550 million penalty under the DSA. AliExpress faces a significant fine: the EU has imposed a record €550 million penalty under the DSA. The EU has imposed a €550 million fine on AliExpress, marking its largest penalty under the DSA so far, for not preventing the sale of counterfeit, unsafe, and illegal items. Vivo X300 FE review: The unexpected compact flagship that impressed me surprisingly. Vivo X300 FE review: The unexpected compact flagship that impressed me surprisingly. The Vivo X300 FE was released at the same time as the X300 Ultra. The FE has been available for a while and might currently be the top compact Android flagship. The 'synthetic insider': AI deepfakes posing as fraudulent employees The 'synthetic insider': AI deepfakes posing as fraudulent employees AI deepfakes enable hackers to impersonate employees, creating the "synthetic insider" threat. However, the majority of insider leaks remain unintentional, and AI agents represent the next potential danger. Images for the Galaxy Watch Ultra 2 have surfaced, supporting speculations regarding its battery, display, and chip. Images for the Galaxy Watch Ultra 2 have surfaced, supporting speculations regarding its battery, display, and chip. Leaked marketing images of the Galaxy Watch Ultra 2 confirm weeks of speculation regarding enhancements in battery, display, and chipset. Leaked marketing images of the Galaxy Watch Ultra 2 support rumors regarding its battery, display, and processor specifications. Leaked marketing images of the Galaxy Watch Ultra 2 support rumors regarding its battery, display, and processor specifications. Leaked marketing visuals of the Galaxy Watch Ultra 2 confirm several weeks' worth of speculation regarding the enhancements to the battery, display, and chipset. A new phishing scam on X is utilizing false login notifications to capture your account information. A new phishing scam on X is utilizing false login notifications to capture your account information. A phishing email that imitates X's genuine login notifications is circulating, aimed at stealing your password rather than safeguarding your account. Here’s how to identify the counterfeit and what steps to take if you inadvertently clicked on it.

Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.

Researchers have showcased GhostWriter, a novel attack method that covertly instills false memories into AI agents, affecting their future responses and independent actions.