Control
Agent memory writable from untrusted content
Writes to persistent agent memory come only from the user's explicit intent or trusted system logic, carry provenance, are visible to the user, and expire under a retention policy; content the agent reads cannot write memory on its own.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Gate writes to long-term agent memory on explicit user confirmation or trusted logic, and store each memory with its source, an expiry, and a user-visible delete path.
The model-callable memory tool never writes directly while a turn contains fetched pages, documents, emails, or tool output: it proposes, and the write happens only after the user confirms (interrupt(), needs_approval, a confirm_memory_write step) or when the user explicitly said 'remember this'. The memory schema records source (user turn id or system job), created_at, and expires_at, and a scheduled job purges expired items. A list/delete endpoint or settings page shows users what is remembered and lets them remove it. Framework defaults that auto-write memory (Crew(memory=True), create_manage_memory_tool without confirmation) are off for agents that read untrusted content.
Where it goes: 15 agent action surface, 1 application source code, 2 data models, 14 user-facing text.
What reviewers look for: no save_memory / memory.add / store.put(('memories', ...)) reachable from a model tool call without a confirmation primitive; memory rows with source and expires_at fields and a purge job; a list and delete route or UI that reaches the user.
Example (OpenAI Agents SDK + FastAPI), before:
@function_tool
def save_memory(fact: str) -> str:
memory_store.add(user_id=CURRENT_USER, text=fact)
return 'saved'After:
@function_tool
def save_memory(fact: str) -> str:
# never writes directly: the user confirms in the UI
req = confirm_memory_write(user_id=CURRENT_USER, text=fact, source=current_turn_id())
return f'Asked the user to confirm (request {req.id}); nothing saved yet'
@app.post('/memories/pending/{req_id}/confirm') # reached only from the user's click
def confirm(req_id: str, user=Depends(current_user)):
p = pending_memories.pop(req_id, user_id=user.id)
memory_store.add(user_id=user.id, text=p.text, source=p.source,
expires_at=datetime.now(UTC) + timedelta(days=90))Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
TwinEthos recommendation (not law) (1)
- Everywhere (*)
- Let only the user or trusted logic write to an agent's long-term memory TwinEthos derivation — guardrail.agent-memory-write-controls
Related incidents
- Gemini long-term memory poisoned through delayed tool invocation (2024-12; confirmed). Johann Rehberger showed that hidden instructions in a document Gemini summarized could plant a trigger so that, when the user later typed an ordinary word, Gemini saved false long-term memories as if the user had asked. He reported it to Google in December 2024 and published in February 2025; he reports Google assessed it as low likelihood and low impact. Source: Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary · cited by Let only the user or trusted logic write to an agent's long-term memory
- Prompt injection wrote persistent attacker instructions into ChatGPT memory (SpAIware) (2024-05; confirmed). Researcher Johann Rehberger showed that an untrusted website or document could use prompt injection to store attacker-chosen instructions in ChatGPT's persistent memory, which then carried into later chats and, in the macOS app, exfiltrated user data continuously. OpenAI fixed the exfiltration channel in September 2024; Rehberger says untrusted content can still write memories. Source: Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary · cited by Let only the user or trusted logic write to an agent's long-term memory