Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.
Large language models are improving their ability to remember user-specific details. Whether it's an individual's writing style, regular tasks, shopping preferences, or project deadlines, AI assistants are increasingly capable of storing long-term memories, making interactions feel more personal and beneficial. However, recent research indicates that this very capability might turn into one of AI’s greatest security weaknesses.
Researchers from New Mexico State University have introduced a new tactic known as GhostWriter, which can secretly insert false memories into AI agents. Instead of outright information theft, this attack manipulates an AI's memory, potentially leading to hazardous decisions long after the initial attack.
This marks a subtle yet crucial change in how AI systems can be compromised. Rather than targeting the AI model directly, attackers focus on its memory.
The attack does not involve hacking the AI; rather, it alters what the AI remembers.
Traditional chatbots typically have minimal or no memory across interactions. On the other hand, contemporary AI agents increasingly depend on persistent memory systems that gather data about users, ongoing projects, and past interactions, enabling them to offer more contextual and personalized replies over time.
The researchers claim that these memory systems introduce a completely new attack surface. GhostWriter operates by subtly injecting harmful information into an AI agent’s long-term memory via concealed prompts or unreliable external content. This false information remains inactive until the AI retrieves it while responding to an otherwise valid request.
For instance, if you ask your AI assistant to summarize emails from your bank, a compromised memory could result in the AI secretly forwarding those emails to an attacker instead. It might also retain incorrect contact details, fake deadlines, wrong preferences, or fabricated information, all due to someone successfully altering the assistant’s perceived truths.
Unlike traditional prompt injection attacks, which typically impact a single conversation, GhostWriter is designed for persistence. Once harmful information enters the memory, it can continue to affect the AI’s behavior across multiple future sessions until it is identified and eliminated.
The researchers outline the attack as a two-phase process. The first phase involves memory injection, where harmful content is silently stored in the AI's memory. The second phase is attack activation, which occurs when the AI unknowingly retrieves that tainted memory while responding to a legitimate user query.
As AI memory becomes increasingly useful, it also becomes an attractive target for attacks.
The timing of this research is notable. Nearly all major AI companies are striving to create assistants that retain user information over weeks, months, or even years. Memory has rapidly turned into one of the industry's key differentiators, as it enables AI to feel less like a basic chatbot and more like a personal assistant.
However, this also means that memory requires protection comparable to the model itself. In their experiments, the researchers found that GhostWriter achieved a memory injection success rate of approximately 98%, with malicious memories being activated about 60% of the time against cutting-edge AI agents. These figures imply that current memory architectures may not yet be able to differentiate between trustworthy information and manipulated data.
The research team is not only addressing the problem but has also proposed a protective framework called Agentic Memory Sentry (AM-Sentry). This framework integrates memory screening with stricter memory management policies. The researchers report that this approach significantly decreased GhostWriter's success rate while maintaining the AI's utility.
As AI agents progress into digital assistants capable of managing emails, scheduling meetings, writing code, and making decisions for users, ensuring the security of their memories may become as crucial as protecting the information they produce. The next challenge in AI security might not just be safeguarding models from harmful prompts, but also defending their memories against being rewritten entirely.
Other articles
Covert prompts can unintentionally alter an AI's memory, and researchers indicate that this poses a significant issue.
Researchers have showcased GhostWriter, a novel attack method that covertly instills false memories into AI agents, affecting their future responses and independent actions.
