TwinEthosRequest access

Standard or framework

NIST SP 800-218A

NIST (U.S. Dept of Commerce) · Everywhere (*) · 2 provisions encoded · verified against the official source as of 2026-10-01.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Official text: nvlpubs.nist.gov.

Standard / soft law

Acquired AI models and components should be verified and scanned before use and loaded safely (NIST SP 800-218A PW.4.4, PW.6.1)

NIST SP 800-218A, PW.4.4 (R1, R2) + PW.6.1 (C1): verify acquired AI models before use; secure model serialization · official text · Soft law or guidance (not binding law)

NIST SP 800-218A asks that any model or model component an organization takes from elsewhere (weights, datasets, reward models, adapters, configuration) be checked for integrity, origin and security, and scanned and tested for vulnerabilities and malicious content, before it is used (PW.4.4), and it points to serialization formats that leave less room for malicious content (PW.6.1). Detect third-party model artifacts loaded without a pinned revision, without a hash or signature check, or through loaders that can execute code (pickle, joblib, torch.load with weights_only=False, trust_remote_code=True).

Who it applies to

  • Duty falls on: developer
  • Organizations that acquire and load third-party AI models or model components (open-weight models, adapters, tokenizers, classical ML model files). Voluntary NIST SSDF community profile.

The guard to add

Pin model downloads to a commit, verify their hash or signature, scan them, and load only safe formats (safetensors, weights_only) without trust_remote_code.

Where the code downloads or loads a model (from_pretrained, hf_hub_download, snapshot_download, torch.load, joblib.load, keras.models.load_model, onnxruntime sessions), the call names an immutable revision (a 40-character commit hash), prefers safetensors (use_safetensors=True, safetensors.torch.load_file), and never enables code execution on load: no trust_remote_code=True for unreviewed repositories, torch.load(..., weights_only=True), keras load_model with safe_mode left on, and no pickle, dill, cloudpickle or joblib load of files that came from outside. Before first use, a helper compares the artifact's SHA-256 with a value committed in the repository (or verifies an OpenSSF model-signing / Sigstore signature), and CI or the ingest job runs a model scanner such as modelscan or picklescan on any pickle-based artifact. Artifacts that fail are rejected, not loaded with a warning.

Where it goes: 1 application source code, 5 dependencies, 8 model configuration, 11 CI/CD pipeline.

What this provision adds:

  • Verify integrity and provenance for every acquired AI component, including datasets, reward models, adaptation layers and configuration parameters, not only model weights.

Example (Hugging Face transformers), before:

model = AutoModelForCausalLM.from_pretrained('acme/support-7b', trust_remote_code=True)

After:

model = AutoModelForCausalLM.from_pretrained(
    'acme/support-7b',
    revision='3f1c2a9e8b7d6c5f4e3a2b1c0d9e8f7a6b5c4d3e',   # pinned, reviewed commit
    use_safetensors=True)                              # no pickle, no remote code

Control: Third-party model artifacts loaded without integrity verification or with a loader that can execute code. The same guard addresses 1 item. Engineering guidance, not legal advice.

Related incidents

No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.

Rule id nist-sp800-218a.pw-4-4-verify-acquired-models · review status: primary source derived

Standard / soft law

AI model inputs and outputs should be validated and encoded so they cannot execute unauthorized code (NIST SP 800-218A PW.5.1)

NIST SP 800-218A, PW.5 / PW.5.1 (R1-R3): secure coding for AI model inputs and outputs · official text · Soft law or guidance (not binding law)

NIST SP 800-218A applies the SSDF secure-coding task PW.5.1 to AI: prompts, user data and model output are all treated as input to be recorded, checked against what the model's context allows, and cleaned or discarded when they fail, and they are encoded so that nothing going into or coming out of a model can run as unauthorized code. Detect model or agent output that reaches an interpreter or renderer (eval or exec, a shell, a SQL query, HTML or Markdown rendering that allows raw HTML or remote images, a file path) without schema validation and encoding or parameterization for that sink.

Who it applies to

  • Duty falls on: developer
  • AI model producers and AI system producers that integrate models into software (the Profile's scope includes incorporating and integrating AI models into other software). Voluntary NIST SSDF community profile.

The guard to add

Treat model output as untrusted input: validate it against a schema, encode or parameterize it for its sink, and run generated code only in a sandbox.

At every point where model or agent output leaves the model call, the code that consumes it applies the control its sink needs. Structured output is parsed with a strict schema (pydantic model_validate_json, zod .parse, or the provider's strict structured-output mode) and rejected, not repaired by another model call, when it does not fit. Chat UIs render Markdown with raw HTML disabled (react-markdown without rehype-raw, or marked output passed through DOMPurify.sanitize) and do not auto-load remote images or link previews from model text (disallowedElements={['img']}, a urlTransform allowlist, or a Content-Security-Policy img-src limited to the app's own origins). Database tools take model values only as bound parameters, shell tools take an argument list with shell=False and an allowlisted executable, file tools resolve paths inside a fixed base directory, and model-written code runs in an isolated sandbox (container or microVM with no credentials and no network by default) instead of eval or exec in the application process. Control characters such as ANSI escape sequences are stripped before output is written to terminals or log viewers.

Where it goes: 9 AI output handling, 1 application source code, 15 agent action surface, 6 API calls and integrations.

What this provision adds:

  • Record and validate model inputs and outputs against what the model's context allows, and clean or discard those that fail.

Example (Next.js chat UI (Vercel AI SDK useChat)), before:

{messages.map(m => (
  <div key={m.id}
    dangerouslySetInnerHTML={{ __html: marked.parse(m.content) }} />
))}

After:

import ReactMarkdown from 'react-markdown';   // escapes raw HTML by default; no rehype-raw
{messages.map(m => (
  <ReactMarkdown key={m.id} disallowedElements={['img']} unwrapDisallowed>
    {m.content}
  </ReactMarkdown>   // no auto-loaded images: model text cannot beacon data out
))}

Control: Model output reaches a code, query, shell, markup, or file-path interpreter without validation or encoding. The same guard addresses 1 item. Engineering guidance, not legal advice.

Related incidents

No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.

Rule id nist-sp800-218a.pw-5-1-model-io-handling · review status: primary source derived