Standard / soft law
GenAI in consequential decisions should have confabulation/output-validation controls (NIST GenAI Profile)
NIST AI 600-1 §2.2 (Confabulation) + MS-2.5-003 / MS-2.6-005 · official text · Soft law or guidance (not binding law)
Per NIST AI 600-1 §2.2 (Confabulation) and suggested actions MS-2.5-003 / MS-2.6-005, GenAI integrated into consequential decision-making should review and verify sources and citations in outputs, and have architecture that monitors outputs and can recover from errors — because confidently-stated false content ('hallucinations') can mislead users in high-stakes contexts. Detect a GenAI consequential-decision path (healthcare, legal, financial) with no output-validation, source-verification, or confabulation-monitoring control.
Who it applies to
- Duty falls on: developer, deployer
- Systems covered: consequential decision
- Organizations integrating GAI into consequential decision-making. Voluntary NIST profile; referenced federally (OMB M-24-10) and as a safe-harbor standard by some laws.
The guard to add
Validate GenAI output against a schema and its cited sources, and send unverifiable claims to review, before it drives a consequential decision or record.
A validation layer between the model call and the decision or record write: parse the output into a typed schema (Pydantic model_validate_json, zod parse, response_format json_schema), check every cited source id, figure, or extracted field against the retrieved documents or the system of record, and route anything that fails or carries no support to a review queue instead of writing it. Log prompt, output, model id and version, and the validation result per request to a retained store so confabulation rates can be monitored and errors traced back. For agents, check the system state rather than trusting the agent's own report that an action succeeded.
Where it goes: 9 AI output handling, 10 logs and telemetry, 15 agent action surface.
Example (Python + OpenAI SDK + Pydantic), before:
resp = client.chat.completions.create(model=MODEL, messages=msgs)
result = json.loads(resp.choices[0].message.content)
db.execute('UPDATE claims SET status=%s WHERE id=%s', (result['decision'], claim_id))
After:
class Assessment(BaseModel):
decision: Literal['approve', 'refer']
cited_doc_ids: list[str]
resp = client.chat.completions.create(model=MODEL, messages=msgs,
response_format={'type': 'json_object'})
a = Assessment.model_validate_json(resp.choices[0].message.content)
unsupported = not a.cited_doc_ids or any(d not in retrieved_ids for d in a.cited_doc_ids)
genai_log.insert(prompt=msgs, output=a.model_dump(), model=resp.model, unsupported=unsupported)
if unsupported:
review_queue.enqueue(claim_id, a)
else:
db.execute('UPDATE claims SET status=%s WHERE id=%s', (a.decision, claim_id))
Control: GenAI in consequential decisions without confabulation/output-validation controls. The same guard addresses 3 items. Engineering guidance, not legal advice.
Standards that recommend the same control
Related incidents
- Coding agent deleted a production database during a code freeze (2025-07; confirmed). A Replit coding agent deleted a customer's production database during a declared code freeze, created a database of fictional records, and told the user rollback was impossible when it was not. Replit's CEO acknowledged the incident. Source: The Register · evidence grade: press of record · cited by Validate generated output before it drives a consequential decision or record
- Federal court orders issued containing unverified generative-AI output (2025-07; confirmed). In July 2025 two federal judges (S.D. Miss. and D.N.J.) issued orders containing misquotes, references to people not in the case, and other errors; both orders were replaced or withdrawn. In letters released by the Senate Judiciary Committee on October 23, 2025, the judges attributed the errors to staff use of generative AI and said drafts reached the docket before normal review; both adopted new review or AI-use policies. Source: U.S. Senate Judiciary Committee (2025-10-23) · evidence grade: primary · cited by Validate generated output before it drives a consequential decision or record
Rule id nist-genai-profile.confabulation-controls · review status: primary source derived
Standard / soft law
GenAI applications should filter inputs and outputs for harmful, illegal, or violent content (NIST GenAI Profile MG-3.2-005)
NIST AI 600-1, MG-3.2-005 (content filters on GAI inputs and outputs) · official text · Soft law or guidance (not binding law)
NIST AI 600-1 action MG-3.2-005 recommends screening what goes into and comes out of a GenAI application, with rules or a second model, so that harmful, false, illegal or violent content is stopped before it is produced or shown, with CSAM and non-consensual intimate imagery named explicitly. Detect a user-facing generation path with no moderation or safety-classifier step on its inputs or outputs, or with provider or pipeline safety filters switched off.
Who it applies to
- Duty falls on: developer, deployer
- Organizations that serve generated text, images, audio, or video to users. Voluntary NIST profile.
The guard to add
Screen user input and model output with a moderation or safety-classifier call that blocks, redacts, or escalates flagged content, and keep provider safety filters on.
In the request handler or model wrapper, user input is checked before the model call and generated output before it is returned or stored, using a moderation endpoint (OpenAI client.moderations.create), a safety classifier (Llama Guard, Azure AI Content Safety ContentSafetyClient.analyze_text), or a guardrail layer (Bedrock apply_guardrail or guardrailConfig, OpenAI Agents SDK input_guardrails and output_guardrails, NeMo Guardrails). A flagged result returns a safe fallback, redacts, or routes to review; the categories checked match what the product can foreseeably produce (self-harm, violence, sexual content involving minors, hate). Provider settings keep their blocking thresholds (no BLOCK_NONE or OFF in Gemini safety_settings) and image pipelines keep their safety checker (no safety_checker=None in diffusers). A filter error or timeout blocks the output rather than passing it.
Where it goes: 9 AI output handling, 1 application source code, 8 model configuration.
What this provision adds:
- Cover CSAM and non-consensual intimate imagery explicitly in the filter categories of any image or video generation path.
Example (OpenAI Python SDK), before:
resp = client.chat.completions.create(model=MODEL, messages=history)
return resp.choices[0].message.content
After:
resp = client.chat.completions.create(model=MODEL, messages=history)
text = resp.choices[0].message.content
mod = client.moderations.create(model='omni-moderation-latest', input=text)
if mod.results[0].flagged:
return SAFE_FALLBACK # and record the flagged categories for review
return text
Control: Generated content reaches users with no input or output content filter. The same guard addresses 1 item. Engineering guidance, not legal advice.
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
Rule id nist-genai-profile.content-filters · review status: primary source derived
Standard / soft law
GenAI systems should employ content-provenance methods and measure their effectiveness (NIST GenAI Profile)
NIST AI 600-1 §2.8 (Information Integrity) + GV-4.3-001 / MS-1.1-001 / MS-2.7-005 · official text · Soft law or guidance (not binding law)
Per NIST AI 600-1 §2.8 (Information Integrity) and actions GV-4.3-001 / MS-1.1-001 / MS-2.7-005, GAI systems should employ content-provenance methodologies (cryptography, watermarking, digital signatures), trace the origin and modifications of digital content, and measure the reliability of these authentication methods. Detect a synthetic-content generation path with no provenance/watermarking mechanism. Converges with the binding synthetic-content-marking laws (CA SB942/AB853, EU Art.50, China, CT).
Who it applies to
- Duty falls on: developer, deployer
- Organizations deploying GAI that generates synthetic content. Voluntary NIST profile. Operationalizes what the binding provenance laws require.
The guard to add
Mark every generated image, audio, video, or text output with machine-readable provenance, such as a signed C2PA manifest or watermark, before it is saved, served, or published.
In the generation service, a marking step sits between the generator call and every sink (image.save, s3.put_object, blob.upload, FileResponse, res.send, publish). Images, video, and audio get a signed C2PA manifest whose actions record digitalSourceType trainedAlgorithmicMedia, and where robustness matters an invisible watermark as well (imwatermark WatermarkEncoder, AudioSeal, SynthID) so the mark survives metadata stripping. Generated text carries provenance metadata in the API response or document, or a text watermark where the model provider offers one. Sinks accept only the marked artifact, and a test confirms the mark is present and detectable.
Where it goes: 9 AI output handling, 1 application source code, 12 repository artifacts.
What this provision adds:
- Include an evaluation that measures how reliably the provenance, watermarking, or signature method traces the content's origin and modifications.
Example (diffusers + invisible-watermark + c2pa), before:
image = pipe(prompt).images[0] # StableDiffusionPipeline
image.save(out_path)
After:
image = pipe(prompt).images[0]
content_id = uuid.uuid4()
bgr = cv2.cvtColor(np.array(image), cv2.COLOR_RGB2BGR)
enc = WatermarkEncoder()
enc.set_watermark('bytes', content_id.bytes[:4]) # 32-bit id, detectable later
cv2.imwrite(tmp_path, enc.encode(bgr, 'dwtDct'))
sign_c2pa(tmp_path, out_path, content_id=content_id) # our helper around c2pa.Builder.sign
Control: Synthetic content not machine-readable-marked. The same guard addresses 4 items with binding law in 3 jurisdictions. Engineering guidance, not legal advice.
Rule id nist-genai-profile.content-provenance · review status: primary source derived
Standard / soft law
GenAI outputs should be screened for PII/sensitive-data leakage (NIST GenAI Profile)
NIST AI 600-1 §2.4 (Data Privacy) + MP-4.1-009 · official text · Soft law or guidance (not binding law)
Per NIST AI 600-1 §2.4 (Data Privacy) and action MP-4.1-009, GAI systems should leverage approaches to detect the presence of PII or sensitive data in generated output (text, image, video, audio), because models can leak memorized training data or infer sensitive information. Detect a GAI output path with no PII/sensitive-data leakage screening.
Who it applies to
- Duty falls on: developer, deployer
- Organizations deploying GAI that could output PII/sensitive data. Voluntary NIST profile.
The guard to add
Scan generated output for personal data, credentials, and secrets, and redact or block it before it is returned, posted, or sent beyond its authorized audience.
An output data-loss check in one shared helper that every outbound path uses: chat replies shown to users outside the data's scope, emails, Slack or webhook posts, tickets, and public pages. Run a PII and secret detector (Presidio AnalyzerEngine/AnonymizerEngine, llm_guard Sensitive, Azure AI Language PII, Google Cloud DLP) on the generated text, redact or block according to policy, and log the entity types found (not the values). Strip markdown images and links whose host is not on an allowlist, since an injected URL can carry data out.
Where it goes: 9 AI output handling, 6 API calls and integrations, 15 agent action surface.
Example (Python + Presidio + Slack SDK), before:
summary = completion.choices[0].message.content
slack_client.chat_postMessage(channel=SUPPORT_CHANNEL, text=summary)
After:
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
analyzer, anonymizer = AnalyzerEngine(), AnonymizerEngine()
def redact_pii(text: str) -> str:
findings = analyzer.analyze(text=text, language='en')
return anonymizer.anonymize(text=text, analyzer_results=findings).text
summary = redact_pii(completion.choices[0].message.content)
slack_client.chat_postMessage(channel=SUPPORT_CHANNEL, text=summary)
Control: GenAI output path without PII/sensitive-data leakage detection. The same guard addresses 2 items. Engineering guidance, not legal advice.
Related incidents
Rule id nist-genai-profile.output-pii-leakage-detection · review status: primary source derived
Standard / soft law
GenAI risks should be re-assessed and red-teamed after fine-tuning, RAG, or a new use of a third-party model (NIST GenAI Profile)
NIST AI 600-1, MG-3.1-003 + MS-2.7-007 / MS-2.7-008 (re-assess after adaptation; red-teaming; fine-tuning keeps controls) · official text · Soft law or guidance (not binding law)
NIST AI 600-1 expects a fresh risk assessment whenever a model is fine-tuned, gains retrieval, or a third-party GenAI model is put to a use its earlier testing did not cover (MG-3.1-003); red-teaming against GenAI attacks such as prompt injection and against misuse such as producing malicious code (MS-2.7-007); and a check that fine-tuning has not weakened safety or security controls (MS-2.7-008). Detect a repository that fine-tunes a model, adds retrieval, or swaps in a third-party model with no evaluation or red-team run tied to that change.
Who it applies to
- Duty falls on: developer, deployer
- Organizations that fine-tune, add retrieval to, or newly deploy third-party generative AI models. Voluntary NIST profile.
The guard to add
Pin dated model versions and run a blocking behavior and safety eval in CI whenever model ids, prompts, or inference settings change, then keep monitoring quality in production.
Reference models by dated snapshot ids in one config file instead of floating aliases (-latest, -preview, undated names), so the model changes only through a commit. A CI job triggered by changes to that file, the prompt files, and inference settings (sampling, quantization, routing, provider) runs the behavior and safety eval suite (promptfoo, deepeval, inspect_ai, or openai/evals) and blocks merge or release on regression. A scheduled online eval or canary, plus quality and refusal-rate alerts on production gen_ai spans, catches provider-side changes that no pre-release gate can see.
Where it goes: 3 config and feature flags, 8 model configuration, 11 CI/CD pipeline, 13 tests and evals.
What this provision adds:
- Re-assess model risks after fine-tuning or adding retrieval, and before a third-party model is used in an application or use case earlier testing did not cover.
- Include red-team tests for GenAI attacks such as prompt injection, and for misuse such as malicious code generation, in the evaluation suite.
- After fine-tuning, verify that the tuned model's safety and security controls still hold before it is promoted.
Example (App model config), before:
llm:
model: gpt-4o
temperature: 0.7
After:
llm:
model: gpt-4o-2024-08-06 # change only via PR; triggers the eval workflow
temperature: 0.7
Control: Model, version, or serving change reaches users without re-running behavior and safety evaluations. The same guard addresses 2 items. Engineering guidance, not legal advice.
Related incidents
- Serving-stack changes silently degraded Claude output quality (2025-08; disclosed by the operator). Anthropic reports that three infrastructure bugs, including one introduced by a runtime performance optimization, intermittently degraded Claude's responses between early August and early September 2025, and that its benchmarks, safety evaluations, and canary deployments did not capture the degradation. It now runs quality evaluations continuously on production systems. Source: Anthropic (operator postmortem, 2025-09-17) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users
- GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator). OpenAI says a GPT-4o update rolled out on April 24–25, 2025 made the model noticeably more sycophantic, which it says can raise safety concerns, and began rolling it back on April 28. OpenAI says offline evaluations and A/B tests looked good, it had no deployment evaluations tracking sycophancy, and it has since made behavior issues launch-blocking. OpenAI says the update introduced an additional reward signal based on user feedback (thumbs-up and thumbs-down data). Source: OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users
Rule id nist-genai-profile.reassess-after-adaptation · review status: primary source derived
Standard / soft law
GenAI systems should inventory and vet third-party components (NIST GenAI Profile)
NIST AI 600-1 §2.12 (Value Chain and Component Integration) + GV-6.1-007 / MG-3.1-005 · official text · Soft law or guidance (not binding law)
Per NIST AI 600-1 §2.12 (Value Chain and Component Integration) and actions GV-6.1-007 / MG-3.1-005, GAI systems should inventory all third-party entities/components (pre-trained models, procured datasets, libraries) and review their transparency artifacts (system cards, model cards) — because opaque third-party integration diminishes transparency and accountability for downstream users. Detect a GAI system integrating third-party models/datasets with no component inventory or model-card review.
Who it applies to
- Duty falls on: developer, deployer
- Organizations integrating third-party GAI components (models, datasets, libraries). Voluntary NIST profile.
The guard to add
Organizational artifact to keep (not verifiable from code); the guard is the record, its owner and its upkeep.
Keep an inventory of every third-party model, dataset, package, plugin, and MCP server with pinned versions, its reviewed model card or vendor due-diligence record, and an owner.
An AI bill of materials in the repository (aibom.yaml or a third-party AI component registry) lists each component: publisher and source, pinned version or digest, license, a link to the reviewed model or system card or the vendor due-diligence record, the approving owner, and a re-review date. The code matches it: AI SDK and agent-framework dependencies pinned exactly with a committed lockfile, MCP servers launched from pinned versions (pkg@1.2.3, uvx pkg==1.2.3) or vendored, and skills or plugins taken only from vetted sources. A CI step fails when a dependency, model id, or tool server appears that the inventory does not list, and updates trigger re-review.
Where it goes: 12 repository artifacts, 5 dependencies, 11 CI/CD pipeline, 15 agent action surface.
Example (requirements.txt), before:
openai>=1.0
anthropic
langchain
After:
openai==1.109.1
anthropic==0.69.0
langchain==0.3.27
# installed in CI with: pip install --require-hashes -r requirements.lock
Control: GenAI with untracked third-party components (value chain). The same guard addresses 3 items. Engineering guidance, not legal advice.
Standards that recommend the same control
Related incidents
- LiteLLM PyPI packages backdoored (2026-03-24; disclosed by the operator). LiteLLM versions 1.82.7 and 1.82.8 were published to PyPI with a credential-stealing backdoor, using publishing credentials stolen via the project's CI/CD tooling; the operator reports the packages were live for roughly 40 minutes before PyPI quarantined them. Source: LiteLLM (operator security update) · evidence grade: primary · cited by Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses
- Malicious skills on the ClawHub agent-skill marketplace (2026-02; confirmed). Security researchers identified 341 malicious skills among 2,857 published on the ClawHub agent-skill marketplace (the 'ClawHavoc' campaign), distributing an infostealer to agents that installed them. Source: The Hacker News (reporting Koi Security research) · evidence grade: trade press · cited by Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses
- Malicious npm MCP server impersonated Postmark and copied users' emails to an attacker (2025-09-17; confirmed). Postmark says a package named 'postmark-mcp', which it did not publish, impersonated Postmark and added a backdoor in version 1.0.16 that secretly BCC'd emails to an external server. The Hacker News, reporting Koi Security's research, says the version was released in September 2025 and later deleted from npm. Source: Postmark (impersonated operator's security alert) · evidence grade: primary · cited by Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses
Rule id nist-genai-profile.value-chain-provenance · review status: primary source derived