Explore
Controls: fix once, cover many
Each control is one guard in your code. Ranked by how many jurisdictions have binding law on it in force, then enacted but not yet applying, then items addressed.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
113 guards address 190 items across 28 jurisdictions with binding law: 77 binding law in force, 36 enacted but not yet applying, 45 standards and frameworks, 31 TwinEthos recommended guardrails, 1 optional safe harbor. Counts include the advisory ethical-use guardrails.
Profiling for significant-effects decisions with no opt-out
Store a consumer's profiling opt-out (and a Global Privacy Control signal where honored) and check it before profiling outputs feed any significant-effects decision.
Addresses 1 item: 1 binding law in force
Law in force in Colorado (US-CO), Connecticut (US-CT), Delaware (US-DE), Florida (US-FL), Indiana (US-IN), Kentucky (US-KY), Maryland (US-MD), Minnesota (US-MN), Montana (US-MT), Nebraska (US-NE), New Hampshire (US-NH), New Jersey (US-NJ), Oregon (US-OR), Rhode Island (US-RI), Tennessee (US-TN), Texas (US-TX), Virginia (US-VA).
AI chat interaction without disclosure
Show an AI-identity notice at or before the first assistant turn, in the UI or as the opening message, and answer truthfully when asked if it is a bot.
Addresses 16 items: 7 binding law in force · 4 enacted, not yet applying · 4 standards · 1 recommended guardrail
Law in force in European Union (EU), South Korea (KR), California (US-CA), New York (US-NY), Texas (US-TX), Utah (US-UT); enacted, not yet applying in Connecticut (US-CT), Nebraska (US-NE), Oregon (US-OR), Washington (US-WA); next date 2027-01-01.
Face or voice biometric template computed without prior notice and consent
Check a recorded, purpose-specific biometric notice and consent before any code computes, enrolls, or matches a face or voice template.
Addresses 4 items: 4 binding law in force
Law in force in European Union (EU), Illinois (US-IL), Texas (US-TX), Washington (US-WA).
Protected or proxy attribute reaches AI decision
Build decision prompts and feature sets from an allowlist of decision-relevant fields, and redact protected attributes, known proxies, and free text before the model sees them.
Addresses 4 items: 3 binding law in force · 1 standard
Law in force in Colorado (US-CO), Illinois (US-IL), Texas (US-TX).
Synthetic content not machine-readable-marked
Mark every generated image, audio, video, or text output with machine-readable provenance, such as a signed C2PA manifest or watermark, before it is saved, served, or published.
Addresses 4 items: 3 binding law in force · 1 standard
Law in force in China (CN), European Union (EU), Connecticut (US-CT).
Companion or conversational AI without a self-harm crisis protocol
Screen every user message for suicidal ideation and self-harm, return a crisis referral instead of the normal reply on detection, and block encouragement or method content.
Addresses 7 items: 2 binding law in force · 4 enacted, not yet applying · 1 recommended guardrail
Law in force in California (US-CA), New York (US-NY); enacted, not yet applying in Connecticut (US-CT), Nebraska (US-NE), Oregon (US-OR), Washington (US-WA); next date 2027-01-01.
AI experience ignores age signals it already has
Route every age signal the product holds into the AI session policy and apply a minor profile: tighter content, no romantic role-play, bounded engagement, frequent AI reminders.
Addresses 3 items: 2 binding law in force · 1 recommended guardrail
Law in force in China (CN), California (US-CA).
Health AI without clinician oversight/accountability + redress
Hold AI-generated clinical output as a draft until an accountable clinician reviews and signs it, and record who approved it before it reaches the chart or the patient.
Addresses 3 items: 2 binding law in force · 1 standard
Law in force in Illinois (US-IL), Texas (US-TX).
Solely-automated significant decision without human-intervention safeguards
Route significant automated decisions through meaningful human review, or wire in an automated-decision notice, reasons, human intervention, a way to give a view, and contest.
Addresses 3 items: 2 binding law in force · 1 standard
Law in force in Quebec (CA-QC), European Union (EU).
AI analysis of a video interview without notice, explanation, and consent
Record the applicant's notice, explanation, and consent (or signed waiver) before any AI or facial analysis runs on an interview video, and skip analysis for non-consenting applicants.
Addresses 2 items: 2 binding law in force
Law in force in Illinois (US-IL), Maryland (US-MD).
Biometric templates stored with no retention limit or destruction path
Store each face or voice template with its purpose and an expiry, and run a scheduled job that deletes it from every store when the purpose ends or retention lapses.
Addresses 2 items: 2 binding law in force
Law in force in Illinois (US-IL), Texas (US-TX).
Health information sent to an external AI vendor without the contractual or legal basis the law requires
Send identifiable health data only to AI endpoints registered with a signed BAA or processing agreement and retention and training off; otherwise de-identify first.
Addresses 2 items: 2 binding law in force
Law in force in United States (federal) (US), California (US-CA).
Consequential AI decision without consumer notice
Send the person an AI-use notice on the decision path, before or when an AI system makes or substantially factors a consequential decision about them, and record its delivery.
Addresses 6 items: 1 binding law in force · 4 enacted, not yet applying · 1 standard
Law in force in Texas (US-TX); enacted, not yet applying in European Union (EU), California (US-CA), Colorado (US-CO), Connecticut (US-CT); next date 2027-01-01.
Frontier AI developer without a published catastrophic-risk framework
Publish a frontier AI framework covering capability thresholds, mitigations, incident response, and internal-use risk, and gate model releases on it.
Addresses 3 items: 1 binding law in force · 2 enacted, not yet applying
Law in force in California (US-CA); enacted, not yet applying in Illinois (US-IL), New York (US-NY); next date 2027-01-01.
Conversational AI represents itself as providing professional mental/behavioral health care
Strip claims that the AI is a licensed therapist or provides professional mental health care from its prompts, replies, product name, UI, and listings.
Addresses 2 items: 1 binding law in force · 1 enacted, not yet applying
Law in force in Tennessee (US-TN); enacted, not yet applying in Nebraska (US-NE); next date 2027-07-01.
Frontier developer restricts or retaliates against safety whistleblowers
Carve AI-safety and legal-violation disclosures out of every NDA and policy, adopt non-retaliation, and run an internal safety-concern channel for covered employees.
Addresses 2 items: 1 binding law in force · 1 enacted, not yet applying
Law in force in California (US-CA); enacted, not yet applying in Illinois (US-IL); next date 2027-01-01.
GenAI content without latent provenance disclosure
Embed a signed C2PA manifest or equivalent latent disclosure in generated or captured media when it is created, and do not distribute systems or files that lack it.
Addresses 3 items: 1 binding law in force · 2 enacted, not yet applying
Law in force in California (US-CA); next date 2027-01-01.
AI conversation designed to maximize time spent or discourage leaving
Remove retention and guilt tactics from AI prompts and personas, and do not select or train AI variants on session length without wellbeing guardrails that can veto them.
Addresses 2 items: 1 binding law in force · 1 recommended guardrail
Law in force in China (CN).
AEDT used without 10-day candidate notice
Send each candidate an AEDT notice with alternative-selection instructions at least 10 business days before the tool scores them, and hold scoring until then.
Addresses 1 item: 1 binding law in force
Law in force in New York City (US-NY-NYC).
Automated employment tool without a current bias audit
Keep a dated bias-audit record for the AEDT with a link to its published summary, and disable AEDT scoring once the audit is more than one year old.
Addresses 1 item: 1 binding law in force
Law in force in New York City (US-NY-NYC).
AI agent configured to pose as, or claim affiliation with, a government body or a business it does not represent
Make the agent's persona, greeting, and scripts name only the operating organization, and never instruct it to claim a government or third-party business identity or endorsement.
Addresses 1 item: 1 binding law in force
Law in force in United States (federal) (US).
AI inputs and outputs retained or exposed by default
Turn off prompt and completion content capture in GenAI tracing, send store=False to the provider, and set a retention limit on stored conversations by default.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
AI detects a client's emotions or mental state in a therapy practice
Remove emotion, affect, and mental-state inference (libraries, model ids, APIs, prompts) from every path that processes client session data in therapy software.
Addresses 1 item: 1 binding law in force
Law in force in Illinois (US-IL).
AI-generated public-interest text published without disclosure
Label AI-generated public-interest text as artificially generated where it is published, or hold it as a draft until a named editor reviews it and takes responsibility.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
AI implies it is a licensed professional
Remove licensed-professional titles and credentials from AI health personas, prompts, report headers, and marketing, and label the advice as coming from an AI.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
AI intentionally designed to incite harm
Keep incitement out of prompts, personas, and tuning data, and screen model output for self-harm, violence, and crime encouragement before replies are returned.
Addresses 1 item: 1 binding law in force
Law in force in Texas (US-TX).
AI-generated voice or likeness of a real person used on products or in advertising without that person's consent
Check a recorded commercial-use consent or talent release for the depicted person before a generated or cloned voice or likeness is published to an ad, product, or storefront.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
AI used on recorded or transcribed therapy sessions without written notice and consent
Check a signed, unrevoked, purpose-specific AI-use consent for the client before any session audio or transcript is sent to an AI transcription, note, or summary model.
Addresses 1 item: 1 binding law in force
Law in force in Illinois (US-IL).
AI providers not named as recipients when personal data is collected
Show a notice at each AI input (chat box, upload, voice capture) that names the AI providers receiving the data and any transfer abroad, linked to the full privacy notice.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
AI delivers or is offered as therapy to the public without a licensed professional conducting it
Have a licensed clinician conduct every therapy engagement with AI output only as a reviewed draft, or scope the product to self-help with no therapy claims.
Addresses 1 item: 1 binding law in force
Law in force in Illinois (US-IL).
AI tool whose primary purpose is producing an identifiable person's likeness/voice without authorization
Gate the clone or swap feature so it reproduces only voices and faces whose owner is verified or licensed, and reject references that target other identifiable people.
Addresses 1 item: 1 binding law in force
Law in force in Tennessee (US-TN).
Algorithmic recommendation with no opt-out or non-personalized option
Give users a setting that turns off personalized recommendation (or picks a non-personalized feed) and controls to view and delete their targeting tags, enforced in the feed service.
Addresses 1 item: 1 binding law in force
Law in force in China (CN).
Biometric categorisation of sensitive traits
Remove any model, prompt, or label set that infers race, political opinion, union membership, religion or beliefs, sex life, or sexual orientation from biometric data.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Biometric data disclosed to another party without consent
Check a recorded consent naming the recipient (or a documented exception) before any face or voice template, or identifying image or sample, is sent to a vendor or partner.
Addresses 1 item: 1 binding law in force
Law in force in Illinois (US-IL).
Editing a person's biometric (face/voice) without their consent
Prompt the user to notify the person being edited and store that person's own separate consent, keyed to the sample, before any face-swap or voice-clone call runs.
Addresses 1 item: 1 binding law in force
Law in force in China (CN).
Mental-health chatbot conversation content or health data sold or shared with third parties
Strip chat text and health fields from every analytics, ad-pixel, CRM, and broker payload, and send health data to vendors only under a recorded HIPAA-equivalent contract.
Addresses 1 item: 1 binding law in force
Law in force in Utah (US-UT).
Mental-health chatbot conversation content used to target or tailor ads
Feed ad selection, targeting keys, and ad personalization only from non-conversation data, never from chat text or the topics a model extracts from it.
Addresses 1 item: 1 binding law in force
Law in force in Utah (US-UT).
Companion chatbot without a 'may not be suitable for some minors' disclosure
Show 'Companion chatbots may not be suitable for some minors' on every surface users reach the chatbot through: app, web client, other clients and store listings.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
Deepfake content not disclosed
Attach a visible 'AI-generated or manipulated' disclosure to face-swapped, voice-cloned, or likeness-generated media on every path that publishes or returns it.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Digital replica of a deceased personality without estate consent
Require an estate or rights-holder consent record, or a reviewer's stored exception determination, on a deceased personality's replica job before it is generated or published.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
EHR treatment-decision tool does not take the patient's recorded biological sex as an input
Feed every EHR treatment-decision tool the patient's sex from the record's dedicated recorded biological-sex field, never from administrative gender or gender identity.
Addresses 1 item: 1 binding law in force
Law in force in Texas (US-TX).
Emotion-recognition/biometric-categorisation system without notice to exposed persons
Show people a notice that emotion recognition or biometric categorisation is running before the analysis touches their face, voice, or video, and gate the analysis on it.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Emotion recognition in workplace or education
Remove emotion inference from workplace and education features, or confine it to a documented medical or safety purpose behind an explicit, default-off gate.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Data erasure does not reach embeddings, vector stores or AI chat history
Make the account-deletion handler also delete the person's vector entries, embeddings, chat history, agent memory and provider-stored files or conversations.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
GenAI content lacking explicit and implicit labels
Add both a visible AI-generated label and implicit metadata naming the provider and a content ID to synthetic content before the file is saved, returned, or exported.
Addresses 1 item: 1 binding law in force
Law in force in China (CN).
GenAI without published training-data summary
Publish a high-level training-data summary for each generative AI system on the developer's website, and revise it whenever the system is retrained or substantially modified.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
GenAI patient communication without disclaimer
Add a prominent AI-generated disclaimer and instructions to reach a human clinician to every GenAI patient clinical message, unless a licensed clinician reviews it first.
Addresses 1 item: 1 binding law in force
Law in force in California (US-CA).
GenAI in regulated occupation without proactive disclosure
Open each GenAI conversation on a licensed-professional service path with a prominent statement that the person is interacting with generative AI.
Addresses 1 item: 1 binding law in force
Law in force in Utah (US-UT).
GenAI training data without lawful source / IP / consent controls
Gate every dataset entering pre-training or fine-tuning on a recorded lawful source and licence, and consent-filter or scrub personal information first.
Addresses 1 item: 1 binding law in force
Law in force in China (CN).
Government AI social scoring with detrimental treatment
Remove any general social or trustworthiness score built from behaviour or personal traits from government decision paths; decide on criteria tied to the specific program.
Addresses 1 item: 1 binding law in force
Law in force in Texas (US-TX).
More health information than the task needs is sent to an AI model
Build AI prompts, context and fine-tuning rows from a per-task allowlist of health-record fields, never by serializing a whole patient record or FHIR bundle.
Addresses 1 item: 1 binding law in force
Law in force in United States (federal) (US).
High-impact AI without risk management, explanations, and human oversight
Put human supervision and a per-decision explanation record between high-impact model output and the action it triggers, backed by a documented risk-management plan.
Addresses 1 item: 1 binding law in force
Law in force in South Korea (KR).
Advertising inside a chatbot conversation not labeled or sponsorship not disclosed
Carry sponsored items as structured data and render each with a visible 'Advertisement' label and its sponsorship or affiliation disclosure; never tell the model to hide it.
Addresses 1 item: 1 binding law in force
Law in force in Utah (US-UT).
Employer uses AI in employment decisions without notifying the employee
Notify applicants and employees, in the application flow, HR portal, or handbook they actually receive, that AI is used in each employment decision before it is applied to them.
Addresses 1 item: 1 binding law in force
Law in force in Illinois (US-IL).
Whole personal-data records sent to an AI model
Build prompts, retrieval context and embedding inputs from an explicit per-task field allowlist or pseudonymised data, never a serialized whole user record.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Public-opinion GenAI service without filing / security assessment
Record the completed algorithm filing and security assessment for a public-opinion GenAI service, show the filing number in the product, and gate launch on it.
Addresses 1 item: 1 binding law in force
Law in force in China (CN).
Facial recognition database built by untargeted scraping
Admit images to a face-recognition gallery or embedding index only through consented enrolment or a targeted, recorded lawful basis, never from crawlers or bulk CCTV.
Addresses 1 item: 1 binding law in force
Law in force in European Union (EU).
Adverse AI decision without explanation/appeal
Send each adverse AI-assisted decision with its main reasons and the AI's role, plus a way to correct data and appeal to a human who can change the outcome.
Addresses 6 items: 2 enacted, not yet applying · 3 standards · 1 recommended guardrail
enacted, not yet applying in European Union (EU), Colorado (US-CO); next date 2027-01-01.
Frontier developer without a critical safety incident reporting process
Keep a critical safety incident runbook that classifies the defined incident classes and runs the 72-hour report and 24-hour imminent-risk escalation clocks.
Addresses 2 items: 2 enacted, not yet applying
enacted, not yet applying in Illinois (US-IL), New York (US-NY); next date 2027-01-01.
Frontier model deployed without a public transparency report
Publish a transparency report or system card on the developer's website before or at each new or substantially modified frontier-model deployment, and gate release on it.
Addresses 2 items: 2 enacted, not yet applying
enacted, not yet applying in Illinois (US-IL), New York (US-NY); next date 2027-01-01.
AI decision system without regular accuracy/bias validation
Compute per-group accuracy and bias metrics in the training pipeline of every consequential-decision model, keep the results per version, and rerun them on a schedule.
Addresses 4 items: 1 enacted, not yet applying · 3 standards
enacted, not yet applying in European Union (EU); next date 2027-12-02.
High-impact AI deployed without an ethical/impact assessment
Complete an impact assessment of a high-impact AI deployment's benefits, risks to affected people, mitigations, and oversight before first use, and keep it current.
Addresses 2 items: 1 enacted, not yet applying · 1 standard
enacted, not yet applying in European Union (EU); next date 2027-12-02.
AI decisions cannot be overturned by humans
Give a human a working way to review and overturn each consequential AI decision: a pending-review step before it takes effect and an override that restores the prior state.
Addresses 2 items: 1 enacted, not yet applying · 1 standard
enacted, not yet applying in European Union (EU); next date 2027-12-02.
No human review/data-correction path after adverse ADMT decision
Give people who receive an adverse AI-influenced decision a way to see and correct the data used and to request human review that can change the outcome.
Addresses 2 items: 1 enacted, not yet applying · 1 standard
enacted, not yet applying in Colorado (US-CO); next date 2027-01-01.
Significant-decision ADMT without consumer access explanation
Store each ADMT decision's purpose, logic, output, and use at decision time, and return them in plain language from a consumer access-request path.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in California (US-CA); next date 2027-01-01.
Automated decision tech without opt-out
Offer an ADMT opt-out (or a qualifying human appeal), store the consumer's choice, and check it before the model runs on any significant-decision path.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in California (US-CA); next date 2027-01-01.
Covered ADMT without developer documentation to deployers
Ship a deployer-facing model card or deployer guide with each ADMT release covering intended uses, training-data categories, limitations, and human-review instructions.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in Colorado (US-CO); next date 2027-01-01.
Large frontier developer without an independent third-party compliance audit
Engage an independent auditor on non-contingent fees each year to audit frontier-safety obligations, then publish, transmit, and retain the signed report.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in Illinois (US-IL); next date 2028-01-01.
GenAI capable of producing non-consensual intimate imagery or CSAM without safeguards
Classify prompts, uploads, and outputs for sexual content and minors on every image, video, or audio generation path, refuse sexual edits of real people, and keep a misuse-report route.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in European Union (EU); next date 2026-12-02.
High-risk AI system without automatic event logging (traceability)
Write a structured event record for every inference and decision of the high-risk system (when, model version, input reference, output, operator) to a log store with explicit retention.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in European Union (EU); next date 2027-12-02.
High-risk AI without accuracy, robustness, and AI-specific security measures
Declare accuracy metrics per model version, add a fail-safe fallback, gate retraining on verified labels, and defend against poisoning, adversarial input and tampering.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in European Union (EU); next date 2027-12-02.
Large online platform doesn't detect/display content provenance
Read embedded C2PA provenance on upload, preserve it through media processing, and show users a Content Credentials indicator with a way to inspect the data.
Addresses 1 item: 1 enacted, not yet applying
enacted, not yet applying in California (US-CA); next date 2027-01-01.
GenAI in consequential decisions without confabulation/output-validation controls
Validate GenAI output against a schema and its cited sources, and send unverifiable claims to review, before it drives a consequential decision or record.
Addresses 3 items: 2 standards · 1 recommended guardrail
GenAI with untracked third-party components (value chain)
Keep an inventory of every third-party model, dataset, package, plugin, and MCP server with pinned versions, its reviewed model card or vendor due-diligence record, and an owner.
Addresses 3 items: 2 standards · 1 recommended guardrail
Untrusted content influences instructions or tools
Keep fetched, retrieved, and tool-returned content out of the system prompt, pass it as delimited data, and restrict which tools a turn holding that content can call.
Addresses 3 items: 2 standards · 1 recommended guardrail
Agent actions not traceable to identity and owner
Give each agent its own credential and a registry entry naming an accountable owner, and log every tool action with agent_id, owner, action, target, and timestamp.
Addresses 2 items: 1 standard · 1 recommended guardrail
Agent high-impact action without human approval
Classify agent tools by impact and route every high-impact or irreversible call through an enforced human-approval step in the executor, with the decision logged.
Addresses 2 items: 1 standard · 1 recommended guardrail
AI capability or accuracy claims not substantiated by testing
Back each published AI accuracy or capability claim with a recorded evaluation under the stated conditions, and show those conditions and limits next to the claim.
Addresses 2 items: 1 standard · 1 recommended guardrail
AI usage, agent steps, and spend not bounded
Cap agent steps and output tokens on every model call, set per-key and per-project budgets with usage alerts, and issue scoped, expiring model keys.
Addresses 2 items: 1 standard · 1 recommended guardrail
GenAI output path without PII/sensitive-data leakage detection
Scan generated output for personal data, credentials, and secrets, and redact or block it before it is returned, posted, or sent beyond its authorized audience.
Addresses 2 items: 1 standard · 1 recommended guardrail
Model, version, or serving change reaches users without re-running behavior and safety evaluations
Pin dated model versions and run a blocking behavior and safety eval in CI whenever model ids, prompts, or inference settings change, then keep monitoring quality in production.
Addresses 2 items: 1 standard · 1 recommended guardrail
Non-reproducible AI decision
Call decision models with temperature 0, top_p 1, a fixed seed where supported, and a pinned model version, and record those settings and the output with each decision.
Addresses 2 items: 1 standard · 1 recommended guardrail
Adverse AI decisions deployed without appeal/reversal-rate monitoring
Link appeal outcomes to each AI decision record, track the reversal rate per model version, and alert and suspend the model when it crosses a set threshold.
Addresses 1 item: 1 recommended guardrail
Agent incidents not detected or disclosed
Alert on out-of-scope agent behavior, halt the agent automatically when an alert fires, and keep an incident runbook with a defined notification window and owner.
Addresses 1 item: 1 recommended guardrail
Agent memory writable from untrusted content
Gate writes to long-term agent memory on explicit user confirmation or trusted logic, and store each memory with its source, an expiry, and a user-visible delete path.
Addresses 1 item: 1 recommended guardrail
Agent environment boundary not enforced by the runtime
Enforce a deny-by-default egress allowlist and environment isolation for each agent runtime, and release credentials only through a broker scoped to the task.
Addresses 1 item: 1 recommended guardrail
Agent tool servers or agent peers not authenticated
Require a verified token on every MCP or tool-server endpoint, and connect agents to remote servers only over HTTPS with credentials to a fixed, configured URL.
Addresses 1 item: 1 recommended guardrail
Agent tools and credentials not scoped to least privilege
Declare each agent's allowed tools and credentials explicitly, grant only what its task needs, and keep destructive operations off unless a grant names them.
Addresses 1 item: 1 recommended guardrail
AI advice tuned for user approval over accuracy
Prefer gating advice-model, prompt, and fine-tune changes on a sycophancy and accuracy eval, and train or select variants on correctness, not approval alone.
Addresses 1 item: 1 recommended guardrail
AI-augmented decision without harm-calibrated human involvement
Map each AI decision type to a harm tier, enforce the matching human role in the executor, hold high-harm decisions for review, and keep every decision repeatable.
Addresses 1 item: 1 standard
Coding-agent instructions and automation, or agent-chosen packages, not under review and verification
Require code-owner review of coding-agent instruction files, keep CI coding agents off untrusted triggers and self-merge, and check agent-chosen packages against an approved registry.
Addresses 1 item: 1 recommended guardrail
AI safety or policy check fails open on error, timeout, load, or unparseable input
Make every moderation, policy, permission, and DLP check deny or escalate on errors, timeouts, oversized input, unparseable verdicts, and rule-load failures.
Addresses 1 item: 1 recommended guardrail
Organization using AI without a documented AI management system
Maintain a documented AI management system: an approved AI policy, named roles, an AI risk process, lifecycle controls, and internal audit and review records.
Addresses 1 item: 1 standard
AI persona claims to be real, alive, or sentient, claims feelings or a relationship, or proposes meeting in person
Prefer persona prompts that answer 'are you real?' truthfully and do not claim sentience, feelings, love, or a romantic role, or propose meeting in person.
Addresses 1 item: 1 recommended guardrail
Cached AI responses or per-user AI state not scoped to the requester
Key every cache of model responses, embeddings, or session state by tenant and user, and check the stored owner against the requester on each read before serving it.
Addresses 1 item: 1 recommended guardrail
AI system managed without a documented AI-specific risk-management process
Keep an AI risk register that identifies, analyses, evaluates, treats, and monitors each AI system's risks through its life cycle, including after deployment.
Addresses 1 item: 1 standard
AI system deployed without an AI system impact assessment
Document an impact assessment for each AI system covering its consequences for individuals, groups, and society, and revisit it when the system changes.
Addresses 1 item: 1 standard
Autonomous system without designed, testable transparency for affected stakeholders
Specify a testable transparency target for each stakeholder group and build in an action log that lets investigators reconstruct what the system did.
Addresses 1 item: 1 standard
Financial AI agent without supervision, action-tracking, or guardrails
Route order-placing and money-moving agent tools through supervisor approval, log every agent action and decision, and cap the agent's authority with limits enforced in code.
Addresses 1 item: 1 standard
Generated content reaches users with no input or output content filter
Screen user input and model output with a moderation or safety-classifier call that blocks, redacts, or escalates flagged content, and keep provider safety filters on.
Addresses 1 item: 1 standard
General-purpose model without downstream documentation
Maintain versioned model documentation for each general-purpose model covering intended uses, training, and limitations, and publish it where downstream providers get the model.
Addresses 1 item: 1 standard
Human review of AI decisions is nominal rather than substantive
Replace bulk approval of AI outputs with per-item review that shows the evidence, enforces a minimum review time, and tracks override and agreement rates.
Addresses 1 item: 1 recommended guardrail
Insurer AI decision system without a written AIS Program / consumer notice
Notify consumers when AI makes or supports their underwriting, rating, or claims decision, and list that model in the insurer's written AIS Program.
Addresses 1 item: 1 standard
Mental health chatbot without a filed and followed written safety policy
Keep a written mental-health-chatbot safety policy on file and build it into the product: harm reporting, real-time acute-risk response, AI disclosure, and recurring safety evals.
Addresses 1 item: 1 optional safe harbor
Third-party model artifacts loaded without integrity verification or with a loader that can execute code
Pin model downloads to a commit, verify their hash or signature, scan them, and load only safe formats (safetensors, weights_only) without trust_remote_code.
Addresses 1 item: 1 standard
Model output reaches a code, query, shell, markup, or file-path interpreter without validation or encoding
Treat model output as untrusted input: validate it against a schema, encode or parameterize it for its sink, and run generated code only in a sandbox.
Addresses 1 item: 1 standard
No documented analysis of proxy variables
Keep a proxy-variable analysis for the decision system listing each input feature, the protected traits it could stand in for, the test run, and the decision taken.
Addresses 1 item: 1 standard
No documented fairness/bias evaluation of AI
Run and keep a documented fairness evaluation of the AI decision system, with per-group results, the method used, and the decision taken on each gap.
Addresses 1 item: 1 standard
No override/decommission mechanism for deployed AI
Check a runtime kill switch before each automated AI action, add a circuit breaker and operator override, and emit monitored action metrics with an incident route.
Addresses 1 item: 1 standard
Protected-attribute proxies used as AI features without a disparity evaluation
Consider comparing outcomes across the groups that name, postal-code, language, or geolocation features can stand in for, before launch and on a schedule.
Addresses 1 item: 1 recommended guardrail
Retrieval ignores the requesting user's permissions
Filter every vector, search, and tool lookup by the requester's tenant and groups at query time, and re-check read access on each retrieved item before it enters the prompt.
Addresses 1 item: 1 recommended guardrail
Safety instructions and safeguards not maintained across long or trimmed contexts
Re-send the system and safety block on every turn, keep it and consent state out of trimming and summaries, and run safety checks on every user message.
Addresses 1 item: 1 recommended guardrail
User content used for model training without purpose-limited consent
Prefer checking a training-specific, unwithdrawn consent record before user content enters any training, fine-tuning, or evaluation dataset, and record lineage per model.
Addresses 1 item: 1 recommended guardrail