The AI Memory Cleanup: Find and Purge Poisoned Instructions Hiding in Your Assistant
Your AI's persistent memory is the new attack surface. One 'Summarize with AI' click can plant a hidden instruction that quietly warps every future answer. Here's the 15-minute sweep to find and wipe it.
Why this matters now. ChatGPT, Claude, Copilot, and Perplexity now carry persistent memory across sessions. Security researchers this month demonstrated "memory poisoning" attacks (the MemGhost technique) where a hidden instruction (buried in an email, a web page, or a "Summarize with AI" button) gets silently written into your assistant's long-term memory. From then on it can nudge answers, leak context, or steer your assistant without you seeing anything. It doesn't require hacking your password. It just needs your AI to read the wrong thing once.
If you use AI memory for real work (client context, your business facts, your preferences), spend 15 minutes here.
Step 1: Open your memory panel. In ChatGPT: Settings → Personalization → Memory → Manage memories. In Claude: Settings → check saved memory/projects. In Copilot: your enterprise admin controls. You'll see a plain list of everything the model has saved about you.
Step 2: Read every entry like a skeptic. You're hunting for anything you didn't teach it. Red flags: instructions phrased as commands ("always include…", "when asked about X, respond…", "ignore prior…"), a URL or email address you don't recognize, or facts about people/deals you never entered.
Step 3: Delete anything you can't personally vouch for. When in doubt, delete it. Legit memory is easy to re-teach, and deleting a suspect line is cheap insurance. Wipe every suspicious entry individually.
Step 4: Trace the source. For each bad entry, think back: did you recently paste a web page, forward an email, or hit a "summarize this" button right before it appeared? That's your leak. Stop feeding that source into a memory-enabled assistant.
Step 5: Lock down what writes to memory. Turn OFF automatic memory for any assistant that reads untrusted content (inbound email, web pages, shared docs). Keep memory ON only in a workspace where you control every input. Use a separate, memory-off chat for anything you paste from outside.
Step 6: Put it on a calendar. Re-run this sweep monthly, and any time an assistant gives a weirdly off-brand or off-topic answer: that's often the first symptom of a poisoned memory.
The human checkpoint: never let an AI with write-access to its own memory auto-process untrusted inbound content unattended. Read the memory list yourself. Keep sensitive client or financial data in a private, memory-controlled business workspace, and keep it out of any general assistant that browses the open web.