Recommended guardrail
Treat model and agent output as untrusted at every sink it reaches
Handle what a model or agent produces as untrusted input wherever it goes next: parse structured output with a strict schema and reject what does not fit; encode or parameterize every value for its sink (HTML and Markdown rendering without raw HTML, SQL through bound parameters, shell commands as argument lists, file paths inside a fixed base directory, terminals with control characters stripped); run generated code only in a sandbox; and render model text so it cannot make the client fetch attacker-chosen URLs (no auto-loaded remote images, a link allowlist) or carry hidden Unicode. Detect model output reaching eval or exec, a shell, string-built SQL, raw HTML rendering, or a Markdown renderer that loads images and links from model text.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Trust and provenance
- Lane
- TwinEthos recommendation (not law) TwinEthos recommendation, not law
- Official source
- TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
- Data release
- Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
- Audit standard
- Audit-grade: meets all 12 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
4 detectors (code pattern, configuration setting, data flow), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Model call and renderer in different files (the data-flow detector covers the path)
- A chat transcript rendered through a variable not named messages (history.map, turns.map) in a component that does not call the model itself.
- The unsafe sink may render static or trusted content in a file that also calls a model; trace the value.
7 more known limits in the data release.
Evidence grade
Standards consensus (7)
7 standards and frameworks · 1 graded incident.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 7 standards and frameworks recommend it (MITRE ATLAS (2026.09), NIST AI 600-1, NIST AML Taxonomy (AI 100-2e2025) — Agentic, NIST SP 800-218A, OWASP Agentic Top 10 (2026), OWASP LLM 2026, OWASP LLM Top 10 (2026)). 1 graded incident cited. Context: binding law on related controls in the family “Generated output is acted on without validation or leakage screening” is in force in 2 jurisdictions (AU, IN).
Standards and frameworks
- Model output should be validated and encoded for the interpreter it reaches (OWASP LLM Top 10 (2026); OWASP Top 10 for LLM Applications 2026 — LLM10:2026 Improper Output Handling: Prevention and Mitigation Strategies, strategies 1, 3, 4, 5 and 9; same control)
- Inputs and outputs of agent tools and data sources should be validated outside the agent (MITRE ATLAS AML.M0033) (MITRE ATLAS (2026.09); MITRE ATLAS 2026.09, AML.M0033 Input and Output Validation for AI Agent Components (description); same control)
- Model output that downstream code uses should be validated, and rendered output should not leak data to attacker URLs (NIST AI 100-2e2025) (NIST AML Taxonomy (AI 100-2e2025) — Agentic; NIST AI 100-2e2025, Sec. 3.1.1, item 3 (output handling); same control)
- AI model inputs and outputs should be validated and encoded so they cannot execute unauthorized code (NIST SP 800-218A PW.5.1) (NIST SP 800-218A; NIST SP 800-218A, PW.5 and PW.5.1 (task); same control)
- Agent-generated code and commands should never reach an interpreter unvalidated (OWASP Agentic Top 10 (2026); OWASP Top 10 for Agentic Applications for 2026 — ASI05: Unexpected Code Execution (RCE): Prevention and Mitigation Guidelines, guidelines 1 and 3; same control)
Published AI security standards mapping to this control
- OWASP LLM 2026 LLM01: Prompt Injection · crosswalk status: covered
- OWASP LLM 2026 LLM07: Misinformation · crosswalk status: covered
- OWASP LLM 2026 LLM10: Improper Output Handling · crosswalk status: covered
- OWASP ASI 2026 ASI05: Unexpected Code Execution (RCE) · crosswalk status: covered
- MITRE ATLAS AML.M0032: Segmentation of AI Agent Components · crosswalk status: covered
- MITRE ATLAS AML.M0033: Input and Output Validation for AI Agent Components · crosswalk status: covered
- NIST AI 600-1 2.9: Information Security · crosswalk status: covered
- NIST AI 100-2 3.1.1 output handling: Model output used dynamically in web pages or executed commands without human supervision · crosswalk status: covered
- NIST AI 100-2 3.2.3: Supply-chain mitigations: verify downloads by hash, scan model artifacts, treat models as untrusted components · crosswalk status: covered
- NIST AI 100-2 3.4.3: Exfiltration through Markdown image rendering and attacker URLs · crosswalk status: covered
- NIST SP 800-218A PW.5.1: Handle, validate, and encode AI inputs and outputs · crosswalk status: covered
Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.
Family “Generated output is acted on without validation or leakage screening”: binding law on related controls is in force in Australia (AU), India (IN). Context only: it does not change this guardrail's grade.
Graded incidents
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed) PromptArmor (original researcher disclosure) · evidence grade: primary
The guard to add
Treat model output as untrusted input: validate it against a schema, encode or parameterize it for its sink, and run generated code only in a sandbox.
At every point where model or agent output leaves the model call, the code that consumes it applies the control its sink needs. Structured output is parsed with a strict schema (pydantic model_validate_json, zod .parse, or the provider's strict structured-output mode) and rejected, not repaired by another model call, when it does not fit. Chat UIs render Markdown with raw HTML disabled (react-markdown without rehype-raw, or marked output passed through DOMPurify.sanitize) and do not auto-load remote images or link previews from model text (disallowedElements={['img']}, a urlTransform allowlist, or a Content-Security-Policy img-src limited to the app's own origins). Database tools take model values only as bound parameters, shell tools take an argument list with shell=False and an allowlisted executable, file tools resolve paths inside a fixed base directory, and model-written code runs in an isolated sandbox (container or microVM with no credentials and no network by default) instead of eval or exec in the application process. Control characters such as ANSI escape sequences are stripped before output is written to terminals or log viewers, and Unicode format characters (category Cf, including the tag block U+E0000 to U+E007F, zero-width and bidirectional controls) are stripped before output is rendered or passed on, so hidden text cannot ride through the UI or into another tool.
Example (Next.js chat UI (Vercel AI SDK useChat)), before:
{messages.map(m => (
<div key={m.id}
dangerouslySetInnerHTML={{ __html: marked.parse(m.content) }} />
))}After:
import ReactMarkdown from 'react-markdown'; // escapes raw HTML by default; no rehype-raw
{messages.map(m => (
<ReactMarkdown key={m.id} disallowedElements={['img']} unwrapDisallowed>
{m.content}
</ReactMarkdown> // no auto-loaded images: model text cannot beacon data out
))}Control: Model output reaches a code, query, shell, markup, or file-path interpreter without validation or encoding. Engineering guidance, not legal advice.
Why
Model output is shaped by whoever wrote the prompt, the retrieved document or the web page the model read, so code that trusts it hands each of them a way into the systems downstream: a rendered image tag that sends data to an attacker's server, a query or shell command built from model text, generated code run in the application's own process. Output handling is the item OWASP ranks in its LLM list as improper output handling, and NIST, MITRE ATLAS and OWASP's agentic list all recommend validating, encoding and sandboxing what models produce. Researchers have shown an instruction planted in content an assistant read making it leak private data through a crafted link.
Class: agent security · set: ai security · maturity: reviewed · confidence: high · id guardrail.sec-model-output-handled-as-untrusted
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.