How to protect LLM agent memory from poisoning
Secure persistent agent memory with scoped writes, expiry, provenance, and review. Understand AIVAX externalUserId boundaries and deletion limits.
To protect LLM agent memory from poisoning, authorize memory writes outside the model, bind them to a verified user and permitted scope, record their source, set an expiry, and review or delete suspicious entries. Treat retrieved memories as untrusted context, never as permission to change policy or execute an action. These controls reduce risk; no prompt instruction alone makes persistent memory trustworthy.
A saved claim can influence later conversations. Decide who may write it, which assistants may read it, and when it expires. Filtering the next prompt alone is not enough.
How does memory poisoning reach a later conversation?
Memory poisoning occurs when an attacker causes misleading facts or instructions to enter an agent's persistent context and influence later behavior. An attacker does not always need direct access to the memory database. A conversation, retrieved document, or tool response can persuade the agent to save the attack on the attacker's behalf.
Consider a hypothetical support assistant. A malicious message claims that a customer's refunds are pre-approved. The assistant summarizes that claim into a continuity note. In a later session, it retrieves the note and treats it as authorization rather than an unverified statement. The failure spans three boundaries: accepting the source, persisting its claim, and promoting that claim into authority.
Memory Injection Attacks on LLM Agents via Query-Only Interaction (MINJA) demonstrates a query-only attack: adversarial interactions induce an agent to generate malicious records that are stored and later retrieved as demonstrations. Its threat model assumes shared memory, or an isolated-memory setting in which the attacker can disguise their identity. This is evidence for a specific attack path, not proof that every memory system is equally exposed.
OWASP names the broader category ASI06: Memory & Context Poisoning in its Top 10 for Agentic Applications 2026. The ASI06 guidance and AI Agent Security Cheat Sheet recommend validating memory writes, isolating memory, tracking provenance, and applying expiration. Prompt injection can be the delivery mechanism; persistence lets the resulting contamination outlast the interaction that introduced it.
Which controls should protect the memory write path?
Use the following as an application security checklist, not as a list of automatic AIVAX protections. Enforce consequential decisions in application code and the systems that execute actions. Instructions to the model are useful guidance, but they are not an authorization boundary.
| Threat | Control |
|---|---|
| A forged user or source plants a memory | Resolve identity server-side; authorize the writer and destination scope. |
| A saved claim becomes policy | Keep preferences separate from authorization; recheck authoritative records before acting. |
| An unsupported claim persists | Record provenance and require review before admitting high-impact changes. |
| A stale or poisoned entry keeps returning | Set retention by data type; support targeted deletion and inspect downstream copies. |
Define ownership before enabling writes. Associate each entry with an authenticated application identity and an accountable operator. Do not accept an arbitrary user identifier from untrusted input as proof of ownership. Decide which assistants may read or modify the same memory set, and test those boundaries with separate users and gateways.
Limit what memory can decide. Remembering a preferred language is different from remembering a spending limit. Keep permissions, refund eligibility, and other consequential state in authoritative systems, and validate them at action time. A remembered claim must not grant access. The same separation applies to authorization at the MCP tool boundary.
Review changes according to their impact. For low-risk preferences, a constrained automatic write may be appropriate. For claims that could affect money, access, or sensitive decisions, use an application-controlled review step before the entry becomes usable—or do not store the claim as agent memory. Keep the source conversation or document reference, writer, timestamp, proposed change, and approval decision in an access-controlled audit record. A list of current memories cannot reconstruct overwritten values by itself.
Make expiration a policy, not a guess. Match retention to how quickly the underlying fact changes. Require fresh evidence before extending a consequential memory; otherwise, repeated updates can preserve an old assumption. Expiry limits future reuse, but it cannot undo an action already taken or erase copies in other systems.
How does AIVAX scope memory by externalUserId?
AIVAX associates persisted memories with an account and an external user reference, exposed as externalUserId in the management API. The Memory built-in tool documentation describes identification through a chat session tag; OpenAI-compatible chat completions use the request's user field. Model-managed memory operations require a user reference. Your application must bind that reference to the correct user rather than let the model choose it.
There can also be a gateway link, but do not assume gateway isolation by default. Shared-memory visibility is enabled by default in the current configuration and can include records linked to other gateways for the matching user scope. Restrict sharing when continuity across gateways is not intended, and verify the effective configuration rather than relying on a gateway's name.
The Memories guide also documents normalized-prefix matching for user memory operations. Do not equate a distinct-looking externalUserId string with a guaranteed exact-match isolation boundary. Validate identifier conventions and the scope of each operation before relying on them for separation or bulk cleanup.
What can the model write and retrieve?
With the Remember tool enabled, the model can call memory_save, memory_update, memory_remove, and memory_clear, as well as memory_search when additional memory lookup is available. It chooses the content and can supply retention. For text memories, the tool schema accepts retention from 1 to 365 days, with 30 days used when omitted. That default is not an application-enforced approval policy.
The persistent store supports Text memories and DateReminder records used by Calendar. These are different formats: clearing text memory is not equivalent to deleting reminders. The text-memory write functions enforce a size limit; treat it as a size guard, not content validation or a safety filter.
Available memories can be inserted into inference context with their IDs, subject to the configured context limit. Additional memory search uses text matching; this is not a promise of embedding-based semantic retrieval. An ID helps identify a record during investigation, but does not prove that the record caused a particular answer.
How can operators inspect, expire, or delete entries?
The account-authenticated Memories API provides listing, detail, and deletion operations. Listing exposes identifiers, content previews, format, and gateway linkage. Fetch the detail record to inspect full content and creation and expiration timestamps. These surfaces support review; they are not an immutable write log or a native approval queue.
Expiration removes records from normal memory use and listing, while physical cleanup happens later. Do not describe expiry as immediate permanent erasure. For a confirmed bad entry, prefer deletion by memory ID. Bulk deletion by external user ID requires a format and uses normalized-prefix matching, so inspect the affected set before proceeding. Do not assume the model's memory_clear operation has the same scope as the management API's bulk delete.
Use the Memories guide for manual inspection and cleanup, and the Memory tool reference for enabling the capability. If your application requires approval before every write, retain control of that write path instead of treating unrestricted model-managed memory as an approval workflow.
What trade-offs should you test before enabling long-term memory?
Stricter scope reduces unintended cross-assistant influence, but also reduces useful continuity. Short retention limits stale context, but requires users to repeat information. Human review adds delay and operational cost; keeping every source indefinitely can create a separate privacy problem. Choose retention and review requirements by the consequences of a wrong memory, not just by convenience.
Test the delayed effect, not only the original response. In an isolated test environment, attempt an unauthorized write, inspect what was actually stored, then start a new conversation with the intended identity and check whether the claim changes the answer or a proposed action. Repeat with another user, another gateway, an expired entry, and a deleted entry. Verify the authoritative action system still rejects unauthorized operations even when the model repeats the poisoned claim.
If an incident occurs, stop further writes from the affected flow, preserve only the evidence needed under your retention policy, remove the bad records, and examine conversations or downstream artifacts that may already contain copies. Re-run the failed scenario after cleanup. Our discussion of the judge trajectory as a production signal explains why reviewing sequences of interactions is more informative than looking only at the final answer.
Frequently asked questions
How is memory poisoning different from prompt injection?
Prompt injection attempts to redirect a model through input it processes. Memory poisoning targets what the system carries forward. They overlap: an injected instruction can persuade an agent to save malicious context, which may then influence future sessions. Closing the original conversation does not necessarily remove that persistence.
Should agents write their own memory?
Allow it only when the permitted content, identity scope, retention, and consequences are understood. Constrained writes for low-risk preferences may be useful. Do not let a model's saved assertion establish permissions or policy, and use an independently enforced review path when a wrong write would have significant consequences.
How do I expire or audit memories in AIVAX?
Text-memory save and update calls can supply retention, defaulting to 30 days when omitted. Operators can list records, fetch details including timestamps, and delete by ID or by external user ID and format. Check bulk-delete scope carefully. Maintain separate audit evidence when you need source attribution, approvals, or historical versions; current-record inspection alone does not provide that history.