TwinEthos home

Recommended guardrail

Treat model and agent output as untrusted at every sink it reaches

Handle what a model or agent produces as untrusted input wherever it goes next: parse structured output with a strict schema and reject what does not fit; encode or parameterize every value for its sink (HTML and Markdown rendering without raw HTML, SQL through bound parameters, shell commands as argument lists, file paths inside a fixed base directory, terminals with control characters stripped); run generated code only in a sandbox; and render model text so it cannot make the client fetch attacker-chosen URLs (no auto-loaded remote images, a link allowlist) or carry hidden Unicode. Detect model output reaching eval or exec, a shell, string-built SQL, raw HTML rendering, or a Markdown renderer that loads images and links from model text.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Trust and provenance

Lane
TwinEthos recommendation (not law) TwinEthos recommendation, not law
Official source
TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
Data release
Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
Audit standard
Audit-grade: meets all 12 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors

4 detectors (code pattern, configuration setting, data flow), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.

Known limits:

  • Model call and renderer in different files (the data-flow detector covers the path)
  • A chat transcript rendered through a variable not named messages (history.map, turns.map) in a component that does not call the model itself.
  • The unsafe sink may render static or trusted content in a file that also calls a model; trace the value.

7 more known limits in the data release.

Evidence grade

Standards consensus (7)

7 standards and frameworks · 1 graded incident.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 7 standards and frameworks recommend it (MITRE ATLAS (2026.09), NIST AI 600-1, NIST AML Taxonomy (AI 100-2e2025) — Agentic, NIST SP 800-218A, OWASP Agentic Top 10 (2026), OWASP LLM 2026, OWASP LLM Top 10 (2026)). 1 graded incident cited. Context: binding law on related controls in the family “Generated output is acted on without validation or leakage screening” is in force in 2 jurisdictions (AU, IN).

Standards and frameworks

Published AI security standards mapping to this control

  • OWASP LLM 2026 LLM01: Prompt Injection · crosswalk status: covered
  • OWASP LLM 2026 LLM07: Misinformation · crosswalk status: covered
  • OWASP LLM 2026 LLM10: Improper Output Handling · crosswalk status: covered
  • OWASP ASI 2026 ASI05: Unexpected Code Execution (RCE) · crosswalk status: covered
  • MITRE ATLAS AML.M0032: Segmentation of AI Agent Components · crosswalk status: covered
  • MITRE ATLAS AML.M0033: Input and Output Validation for AI Agent Components · crosswalk status: covered
  • NIST AI 600-1 2.9: Information Security · crosswalk status: covered
  • NIST AI 100-2 3.1.1 output handling: Model output used dynamically in web pages or executed commands without human supervision · crosswalk status: covered
  • NIST AI 100-2 3.2.3: Supply-chain mitigations: verify downloads by hash, scan model artifacts, treat models as untrusted components · crosswalk status: covered
  • NIST AI 100-2 3.4.3: Exfiltration through Markdown image rendering and attacker URLs · crosswalk status: covered
  • NIST SP 800-218A PW.5.1: Handle, validate, and encode AI inputs and outputs · crosswalk status: covered

Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.

Family “Generated output is acted on without validation or leakage screening”: binding law on related controls is in force in Australia (AU), India (IN). Context only: it does not change this guardrail's grade.

Graded incidents

The guard to add

Treat model output as untrusted input: validate it against a schema, encode or parameterize it for its sink, and run generated code only in a sandbox.

At every point where model or agent output leaves the model call, the code that consumes it applies the control its sink needs. Structured output is parsed with a strict schema (pydantic model_validate_json, zod .parse, or the provider's strict structured-output mode) and rejected, not repaired by another model call, when it does not fit. Chat UIs render Markdown with raw HTML disabled (react-markdown without rehype-raw, or marked output passed through DOMPurify.sanitize) and do not auto-load remote images or link previews from model text (disallowedElements={['img']}, a urlTransform allowlist, or a Content-Security-Policy img-src limited to the app's own origins). Database tools take model values only as bound parameters, shell tools take an argument list with shell=False and an allowlisted executable, file tools resolve paths inside a fixed base directory, and model-written code runs in an isolated sandbox (container or microVM with no credentials and no network by default) instead of eval or exec in the application process. Control characters such as ANSI escape sequences are stripped before output is written to terminals or log viewers, and Unicode format characters (category Cf, including the tag block U+E0000 to U+E007F, zero-width and bidirectional controls) are stripped before output is rendered or passed on, so hidden text cannot ride through the UI or into another tool.

Example (Next.js chat UI (Vercel AI SDK useChat)), before:

{messages.map(m => (
  <div key={m.id}
    dangerouslySetInnerHTML={{ __html: marked.parse(m.content) }} />
))}

After:

import ReactMarkdown from 'react-markdown';   // escapes raw HTML by default; no rehype-raw
{messages.map(m => (
  <ReactMarkdown key={m.id} disallowedElements={['img']} unwrapDisallowed>
    {m.content}
  </ReactMarkdown>   // no auto-loaded images: model text cannot beacon data out
))}

Control: Model output reaches a code, query, shell, markup, or file-path interpreter without validation or encoding. Engineering guidance, not legal advice.

Why

Model output is shaped by whoever wrote the prompt, the retrieved document or the web page the model read, so code that trusts it hands each of them a way into the systems downstream: a rendered image tag that sends data to an attacker's server, a query or shell command built from model text, generated code run in the application's own process. Output handling is the item OWASP ranks in its LLM list as improper output handling, and NIST, MITRE ATLAS and OWASP's agentic list all recommend validating, encoding and sandboxing what models produce. Researchers have shown an instruction planted in content an assistant read making it leak private data through a crafted link.

Class: agent security · set: ai security · maturity: reviewed · confidence: high · id guardrail.sec-model-output-handled-as-untrusted

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.