Recommended guardrail
Sanitize and label the source and trust tier of everything written to an index or agent memory
Before content from the web, uploads, messages or tools is embedded into a retrieval index or written to agent memory, normalize it and strip invisible characters, attach its source, ingestion time and trust tier, and keep trust tiers in separate collections so retrieval can filter and weight by trust. Detect ingest and memory-write paths that take external content with no sanitization and no trust-tier label. Gating of who may write memory is covered by guardrail.agent-memory-write-controls; isolating untrusted content in the prompt by guardrail.agent-untrusted-content-isolation.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Trust and provenance
- Lane
- TwinEthos recommendation (not law) TwinEthos recommendation, not law
- Official source
- TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
- Data release
- Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
- Audit standard
- Audit-grade: meets all 12 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
1 detector (data flow), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Sanitization or labelling done in another module before the write is not seen; confirm the call path.
- Loader metadata such as 'source' alone is not counted as a trust tier.
- Content that reaches the write through the model (an agent tool that saves what the model chose) spans two functions; guardrail.agent-memory-write-controls covers that path.
Evidence grade
Standards consensus (5)
5 standards and frameworks · 3 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 5 standards and frameworks recommend it (MITRE ATLAS, NIST AI 100-2, OWASP ACS, OWASP ASI 2026, OWASP LLM 2026). 3 graded incidents cited.
Published AI security standards mapping to this control
- OWASP LLM 2026 LLM01: Prompt Injection · crosswalk status: covered
- OWASP LLM 2026 LLM05: Data and Model Poisoning · crosswalk status: covered
- OWASP LLM 2026 LLM09: Vector and Embedding Weaknesses · crosswalk status: partial
- OWASP ASI 2026 ASI01: Agent Goal Hijack · crosswalk status: partial
- OWASP ASI 2026 ASI06: Memory and Context Poisoning · crosswalk status: covered
- OWASP ACS ACS-Provenance: Field-level provenance (origin, source, derived_from) and a trust label that never rises on derived data · crosswalk status: partial
- MITRE ATLAS AML.M0025: Maintain AI Dataset Provenance · crosswalk status: covered
- MITRE ATLAS AML.M0031: Memory Hardening · crosswalk status: covered
- NIST AI 100-2 3.4.4: Indirect prompt injection: filter third-party instructions, spotlighting, differently-permissioned models, well-defined interfaces · crosswalk status: covered
Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Gemini long-term memory poisoned through delayed tool invocation (2024-12; confirmed) Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed) PromptArmor (original researcher disclosure) · evidence grade: primary
- Prompt injection wrote persistent attacker instructions into ChatGPT memory (SpAIware) (2024-05; confirmed) Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary
The guard to add
Sanitize content and label its source and trust tier before it enters an index or memory; keep trust tiers apart.
In every ingest job and memory writer, before add_documents, from_documents, VectorStoreIndex.from_documents, collection.add, upsert, Memory.add or store.put: normalize text (unicodedata.normalize('NFKC', ...)), strip zero-width (U+200B to U+200D, U+2060), tag-block (U+E0000 to U+E007F) and variation-selector (U+FE00 to U+FE0F) characters, and attach metadata with the source, the ingestion time and a trust tier (operator, verified partner, user-supplied, public web). Write each trust tier to its own collection or namespace, never feed the agent's own outputs back into trusted memory, and let retrieval filter or down-weight by trust tier so untrusted text cannot outrank the operator's own.
Example (Python + LangChain + Chroma), before:
docs = WebBaseLoader(url).load()
vectorstore.add_documents(docs)After:
INVISIBLE = re.compile('[\u200b-\u200d\u2060\ufe00-\ufe0f\U000e0000-\U000e007f]')
docs = WebBaseLoader(url).load()
for d in docs:
d.page_content = INVISIBLE.sub('', unicodedata.normalize('NFKC', d.page_content))
d.metadata.update(source=url, ingested_at=now_iso(), trust_tier='public_web')
public_web_store.add_documents(docs) # its own collection, never the operator'sControl: Content enters a retrieval index or agent memory without provenance, trust tier or sanitization. Engineering guidance, not legal advice.
Why
Retrieval indexes and agent memory turn one poisoned page into a standing instruction: researchers have shown instructions planted in a public channel, a web page or a document being retrieved into an answer or written into memory and carried into later sessions. Labelling where content came from, stripping the invisible characters used to hide instructions, and keeping untrusted sources apart is cheap at ingest and very hard to retrofit once an index is mixed. OWASP recommends it, and MITRE ATLAS and NIST recommend related controls.
Class: agent security · set: agent containment · maturity: reviewed · confidence: high · id guardrail.agent-ingest-provenance-and-sanitization
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.