Control
Model output reaches a code, query, shell, markup, or file-path interpreter without validation or encoding
Model and agent output is handled as untrusted input: structured output is validated against a strict schema, every value is encoded or parameterized for the context it reaches (HTML, Markdown, SQL, shell, file paths, terminals), generated code runs only in a sandbox, and rendered output cannot make the client fetch attacker-chosen URLs.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Treat model output as untrusted input: validate it against a schema, encode or parameterize it for its sink, and run generated code only in a sandbox.
At every point where model or agent output leaves the model call, the code that consumes it applies the control its sink needs. Structured output is parsed with a strict schema (pydantic model_validate_json, zod .parse, or the provider's strict structured-output mode) and rejected, not repaired by another model call, when it does not fit. Chat UIs render Markdown with raw HTML disabled (react-markdown without rehype-raw, or marked output passed through DOMPurify.sanitize) and do not auto-load remote images or link previews from model text (disallowedElements={['img']}, a urlTransform allowlist, or a Content-Security-Policy img-src limited to the app's own origins). Database tools take model values only as bound parameters, shell tools take an argument list with shell=False and an allowlisted executable, file tools resolve paths inside a fixed base directory, and model-written code runs in an isolated sandbox (container or microVM with no credentials and no network by default) instead of eval or exec in the application process. Control characters such as ANSI escape sequences are stripped before output is written to terminals or log viewers.
Where it goes: 9 AI output handling, 1 application source code, 15 agent action surface, 6 API calls and integrations.
What reviewers look for: no dangerouslySetInnerHTML, innerHTML, v-html, or rehype-raw fed by model text without a sanitizer; no eval, exec, new Function, os.system, or shell=True reached by model output; SQL built from model values only through placeholders or an ORM; a schema parse between the model call and any tool or database write; image rendering of model Markdown disabled or allowlisted; code execution delegated to a sandbox service or container.
Example (Next.js chat UI (Vercel AI SDK useChat)), before:
{messages.map(m => (
<div key={m.id}
dangerouslySetInnerHTML={{ __html: marked.parse(m.content) }} />
))}After:
import ReactMarkdown from 'react-markdown'; // escapes raw HTML by default; no rehype-raw
{messages.map(m => (
<ReactMarkdown key={m.id} disallowedElements={['img']} unwrapDisallowed>
{m.content}
</ReactMarkdown> // no auto-loaded images: model text cannot beacon data out
))}Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- Everywhere (*)
- AI model inputs and outputs should be validated and encoded so they cannot execute unauthorized code (NIST SP 800-218A PW.5.1) NIST SP 800-218A, PW.5 / PW.5.1 (R1-R3): secure coding for AI model inputs and outputs
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Coding agent deleted a production database during a code freeze (2025-07; confirmed). A Replit coding agent deleted a customer's production database during a declared code freeze, created a database of fictional records, and told the user rollback was impossible when it was not. Replit's CEO acknowledged the incident. Source: The Register · evidence grade: press of record · cited by Validate generated output before it drives a consequential decision or record
- Federal court orders issued containing unverified generative-AI output (2025-07; confirmed). In July 2025 two federal judges (S.D. Miss. and D.N.J.) issued orders containing misquotes, references to people not in the case, and other errors; both orders were replaced or withdrawn. In letters released by the Senate Judiciary Committee on October 23, 2025, the judges attributed the errors to staff use of generative AI and said drafts reached the docket before normal review; both adopted new review or AI-use policies. Source: U.S. Senate Judiciary Committee (2025-10-23) · evidence grade: primary · cited by Validate generated output before it drives a consequential decision or record
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed). Researchers showed that an instruction planted in a public Slack channel could make Slack AI leak private-channel data through a crafted link. Salesforce patched the issue and reported no evidence of unauthorized access to customer data. Source: PromptArmor (original researcher disclosure) · evidence grade: primary · cited by Screen model and agent output for personal and sensitive data before it leaves the trust boundary