Recommended guardrail
Let only the user or trusted logic write to an agent's long-term memory
Gate writes to persistent agent memory on explicit user intent or trusted system logic, record where each memory came from, show memories to the user with a way to delete them, and expire them under a retention policy. Detect memory-write tools callable by the model while it processes untrusted content, with no confirmation step. Injection into the current turn is covered by guardrail.agent-untrusted-content-isolation; this rule covers what persists.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Standards consensus (2)
2 standards and frameworks · 2 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (NIST AML Taxonomy (AI 100-2e2025) — Agentic, OWASP LLM Top 10 (2025)). 2 graded incidents cited.
Standards and frameworks
- AI agents should mitigate indirect prompt injection and agent hijacking (NIST AI 100-2e2025) (NIST AML Taxonomy (AI 100-2e2025) — Agentic; NIST AI 100-2e2025, Secs. 3.4 (Indirect Prompt Injection Attacks and Mitigations), 3.5 (Security of Agents) and 3.6 (Benchmarks for AML Vulnerabilities); cited)
- Untrusted external content should not flow into agent instructions or tool calls without mediation (OWASP LLM Top 10 (2025); OWASP Top 10 for LLM Applications (2025) — LLM01: Prompt Injection; cited)
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Gemini long-term memory poisoned through delayed tool invocation (2024-12; confirmed) Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary
- Prompt injection wrote persistent attacker instructions into ChatGPT memory (SpAIware) (2024-05; confirmed) Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary
The guard to add
Gate writes to long-term agent memory on explicit user confirmation or trusted logic, and store each memory with its source, an expiry, and a user-visible delete path.
The model-callable memory tool never writes directly while a turn contains fetched pages, documents, emails, or tool output: it proposes, and the write happens only after the user confirms (interrupt(), needs_approval, a confirm_memory_write step) or when the user explicitly said 'remember this'. The memory schema records source (user turn id or system job), created_at, and expires_at, and a scheduled job purges expired items. A list/delete endpoint or settings page shows users what is remembered and lets them remove it. Framework defaults that auto-write memory (Crew(memory=True), create_manage_memory_tool without confirmation) are off for agents that read untrusted content.
Example (OpenAI Agents SDK + FastAPI), before:
@function_tool
def save_memory(fact: str) -> str:
memory_store.add(user_id=CURRENT_USER, text=fact)
return 'saved'After:
@function_tool
def save_memory(fact: str) -> str:
# never writes directly: the user confirms in the UI
req = confirm_memory_write(user_id=CURRENT_USER, text=fact, source=current_turn_id())
return f'Asked the user to confirm (request {req.id}); nothing saved yet'
@app.post('/memories/pending/{req_id}/confirm') # reached only from the user's click
def confirm(req_id: str, user=Depends(current_user)):
p = pending_memories.pop(req_id, user_id=user.id)
memory_store.add(user_id=user.id, text=p.text, source=p.source,
expires_at=datetime.now(UTC) + timedelta(days=90))Control: Agent memory writable from untrusted content. Engineering guidance, not legal advice.
Why
Persistent memory can make a single injection durable: planted instructions or false facts carry into later sessions. A security researcher has shown this in more than one assistant, using documents and web pages the user only asked the assistant to read.
Class: agent security · set: agent containment · maturity: reviewed · confidence: high · id guardrail.agent-memory-write-controls