TwinEthosRequest access

Explore

Recommended guardrails, graded by evidence

What a responsible AI integration does anyway, and how far law, standards and real incidents already back each practice.

TwinEthos recommendations — not law

Guardrails are TwinEthos's opinion, never law. The grade is computed from the corpus: law in force on the same control (or in provisions a guardrail cites as convergence), law enacted but not yet applying, standards that recommend it, and graded incidents. Ethical-use guardrails are advisory.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

31 guardrails. Law in force: 10 · Law coming: 8 · Standards: 11 · Ahead of the law: 2.

Law derived (8)

GuardrailEvidence gradeEvidence
Explain adverse AI-assisted decisions and offer a way to contest them — everywhereLaw in force in 1 jurisdiction, coming in 2 moreLaw in force in 1 jurisdiction · law coming in 2 more · 3 standards and frameworks · 1 graded incident
Tell people when they are interacting with AI — everywhere, not only where requiredLaw in force in 6 jurisdictions, coming in 4 moreLaw in force in 6 jurisdictions · law coming in 4 more · 4 standards and frameworks · 1 graded incident
Run a self-harm crisis protocol in any conversational AI that users may confide inLaw in force in 2 jurisdictions, coming in 4 moreLaw in force in 2 jurisdictions · law coming in 4 more · 2 graded incidents
Validate generated output before it drives a consequential decision or recordStandards consensus (2)2 standards and frameworks · 2 graded incidents
Record enough at decision time to reproduce and explain every consequential AI decisionLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 1 standard or framework · 1 graded incident
Screen model and agent output for personal and sensitive data before it leaves the trust boundaryRecommended by 1 standard1 standard or framework · 1 graded incident
Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrongStandards consensus (2)2 standards and frameworks · 1 graded incident
Make human review of adverse AI decisions substantive, not nominalLaw in force in 1 jurisdiction, coming in 1 moreLaw in force in 1 jurisdiction · law coming in 1 more · 1 standard or framework · 2 graded incidents

Agent security (11)

GuardrailEvidence gradeEvidence
Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing themStandards consensus (2)2 standards and frameworks · 2 graded incidents
Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent usesStandards consensus (2)2 standards and frameworks · 3 graded incidents
Enforce an agent's environment boundary in the runtime, not in the agent's judgmentLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 1 standard or framework · 4 graded incidents
Require human approval before an agent takes a high-impact or irreversible actionLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident
Give every agent a verifiable identity bound to an accountable owner, and log every actionLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 1 standard or framework · 2 graded incidents
Detect out-of-scope agent behavior, halt it, and disclose incidents on a defined clockLaw in force in 1 jurisdiction, coming in 1 moreLaw in force in 1 jurisdiction · law coming in 1 more · 3 graded incidents
Scope every agent's tools and credentials to least privilegeStandards consensus (2)2 standards and frameworks · 2 graded incidents
Let only the user or trusted logic write to an agent's long-term memoryStandards consensus (2)2 standards and frameworks · 2 graded incidents
Enforce the requesting user's permissions on every retrievalStandards consensus (3)3 standards and frameworks · 2 graded incidents
Authenticate every agent tool server and verify every server an agent connects toStandards consensus (2)2 standards and frameworks · 2 graded incidents
Segregate untrusted content from an agent's instructions and tool invocationsLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident

Operational integrity (5)

GuardrailEvidence gradeEvidence
Bound every AI workload's steps, tokens, and spend, and alert on anomaliesRecommended by 1 standard1 standard or framework · 1 graded incident
Make AI safety and policy checks fail closed on error, timeout, load, or unparseable inputLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 3 graded incidents
Re-run behavior and safety evaluations before any model, version, or serving change reaches usersLaw coming in 1 jurisdictionLaw coming in 1 jurisdiction · 2 standards and frameworks · 2 graded incidents
Keep safety instructions and safeguards in force for the whole conversationLaw coming in 3 jurisdictionsLaw coming in 3 jurisdictions · 2 graded incidents
Scope every AI response and session cache to the requesting user or tenantAhead of the law: no law yet, 2 incidents2 graded incidents

Ethical use (advisory, opt-in) (7)

GuardrailEvidence gradeEvidence
Evaluate advice-giving AI for sycophancy, and do not tune it on approval aloneAhead of the law: no law yet, 2 incidents2 graded incidents
Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traitsLaw in force in 2 jurisdictionsLaw in force in 2 jurisdictions · 3 standards and frameworks · 4 graded incidents
Apply minor-appropriate AI settings whenever the product already has an age signalLaw in force in 2 jurisdictions, coming in 2 moreLaw in force in 2 jurisdictions · law coming in 2 more · 2 graded incidents
Keep AI personas from claiming feelings, a real existence, or a relationship, and from proposing to meetLaw in force in 1 jurisdiction, coming in 3 moreLaw in force in 1 jurisdiction · law coming in 3 more · 2 graded incidents
Do not design AI conversations to maximize time spent or to discourage leavingLaw in force in 1 jurisdiction, coming in 1 moreLaw in force in 1 jurisdiction · law coming in 1 more · 2 graded incidents
Train on user content only with consent specific to that purposeLaw in force in 1 jurisdictionLaw in force in 1 jurisdiction · 1 graded incident
Back every AI capability or accuracy claim with testing under the conditions it describesRecommended by 1 standard1 standard or framework · 2 graded incidents