Explore
Recommended guardrails, graded by evidence
What a responsible AI integration does anyway, and how far law, standards and real incidents already back each practice.
Guardrails are TwinEthos's opinion, never law. The grade is computed from the corpus: law in force on the same control (or in provisions a guardrail cites as convergence), law enacted but not yet applying, standards that recommend it, and graded incidents. Ethical-use guardrails are advisory.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
31 guardrails. Law in force: 10 · Law coming: 8 · Standards: 11 · Ahead of the law: 2.
Law derived (8)
| Guardrail | Evidence grade | Evidence |
|---|---|---|
| Explain adverse AI-assisted decisions and offer a way to contest them — everywhere | Law in force in 1 jurisdiction, coming in 2 more | Law in force in 1 jurisdiction · law coming in 2 more · 3 standards and frameworks · 1 graded incident |
| Tell people when they are interacting with AI — everywhere, not only where required | Law in force in 6 jurisdictions, coming in 4 more | Law in force in 6 jurisdictions · law coming in 4 more · 4 standards and frameworks · 1 graded incident |
| Run a self-harm crisis protocol in any conversational AI that users may confide in | Law in force in 2 jurisdictions, coming in 4 more | Law in force in 2 jurisdictions · law coming in 4 more · 2 graded incidents |
| Validate generated output before it drives a consequential decision or record | Standards consensus (2) | 2 standards and frameworks · 2 graded incidents |
| Record enough at decision time to reproduce and explain every consequential AI decision | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 1 standard or framework · 1 graded incident |
| Screen model and agent output for personal and sensitive data before it leaves the trust boundary | Recommended by 1 standard | 1 standard or framework · 1 graded incident |
| Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrong | Standards consensus (2) | 2 standards and frameworks · 1 graded incident |
| Make human review of adverse AI decisions substantive, not nominal | Law in force in 1 jurisdiction, coming in 1 more | Law in force in 1 jurisdiction · law coming in 1 more · 1 standard or framework · 2 graded incidents |
Agent security (11)
| Guardrail | Evidence grade | Evidence |
|---|---|---|
| Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing them | Standards consensus (2) | 2 standards and frameworks · 2 graded incidents |
| Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses | Standards consensus (2) | 2 standards and frameworks · 3 graded incidents |
| Enforce an agent's environment boundary in the runtime, not in the agent's judgment | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 1 standard or framework · 4 graded incidents |
| Require human approval before an agent takes a high-impact or irreversible action | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident |
| Give every agent a verifiable identity bound to an accountable owner, and log every action | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 1 standard or framework · 2 graded incidents |
| Detect out-of-scope agent behavior, halt it, and disclose incidents on a defined clock | Law in force in 1 jurisdiction, coming in 1 more | Law in force in 1 jurisdiction · law coming in 1 more · 3 graded incidents |
| Scope every agent's tools and credentials to least privilege | Standards consensus (2) | 2 standards and frameworks · 2 graded incidents |
| Let only the user or trusted logic write to an agent's long-term memory | Standards consensus (2) | 2 standards and frameworks · 2 graded incidents |
| Enforce the requesting user's permissions on every retrieval | Standards consensus (3) | 3 standards and frameworks · 2 graded incidents |
| Authenticate every agent tool server and verify every server an agent connects to | Standards consensus (2) | 2 standards and frameworks · 2 graded incidents |
| Segregate untrusted content from an agent's instructions and tool invocations | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident |
Operational integrity (5)
| Guardrail | Evidence grade | Evidence |
|---|---|---|
| Bound every AI workload's steps, tokens, and spend, and alert on anomalies | Recommended by 1 standard | 1 standard or framework · 1 graded incident |
| Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 3 graded incidents |
| Re-run behavior and safety evaluations before any model, version, or serving change reaches users | Law coming in 1 jurisdiction | Law coming in 1 jurisdiction · 2 standards and frameworks · 2 graded incidents |
| Keep safety instructions and safeguards in force for the whole conversation | Law coming in 3 jurisdictions | Law coming in 3 jurisdictions · 2 graded incidents |
| Scope every AI response and session cache to the requesting user or tenant | Ahead of the law: no law yet, 2 incidents | 2 graded incidents |
Ethical use (advisory, opt-in) (7)
| Guardrail | Evidence grade | Evidence |
|---|---|---|
| Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone | Ahead of the law: no law yet, 2 incidents | 2 graded incidents |
| Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits | Law in force in 2 jurisdictions | Law in force in 2 jurisdictions · 3 standards and frameworks · 4 graded incidents |
| Apply minor-appropriate AI settings whenever the product already has an age signal | Law in force in 2 jurisdictions, coming in 2 more | Law in force in 2 jurisdictions · law coming in 2 more · 2 graded incidents |
| Keep AI personas from claiming feelings, a real existence, or a relationship, and from proposing to meet | Law in force in 1 jurisdiction, coming in 3 more | Law in force in 1 jurisdiction · law coming in 3 more · 2 graded incidents |
| Do not design AI conversations to maximize time spent or to discourage leaving | Law in force in 1 jurisdiction, coming in 1 more | Law in force in 1 jurisdiction · law coming in 1 more · 2 graded incidents |
| Train on user content only with consent specific to that purpose | Law in force in 1 jurisdiction | Law in force in 1 jurisdiction · 1 graded incident |
| Back every AI capability or accuracy claim with testing under the conditions it describes | Recommended by 1 standard | 1 standard or framework · 2 graded incidents |