Control
AI-augmented decision without harm-calibrated human involvement
The degree of human involvement in AI-augmented decisions (in/over/out-of-the-loop) should be calibrated to the severity and probability of harm; and the decision process should be explainable or at least repeatable.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Map each AI decision type to a harm tier, enforce the matching human role in the executor, hold high-harm decisions for review, and keep every decision repeatable.
An oversight matrix in config (decision type to harm severity and probability, and to human-in-the-loop, human-over-the-loop, or human-out-of-the-loop) that the decision executor reads before any outcome takes effect. High-harm decisions are written with status='pending_review' (or paused with interrupt_before / can_use_tool on an agent) and only a reviewer can confirm, amend, or reverse them; over-the-loop decisions apply but surface to a supervisor who can intervene; only low-harm decisions run unattended. Each decision stores what is needed to explain or repeat it: pinned model version, temperature=0 or a fixed seed, logged inputs and outputs, and reason codes or feature attributions.
Where it goes: 3 config and feature flags, 1 application source code, 15 agent action surface, 10 logs and telemetry.
What reviewers look for: a written harm-to-involvement mapping that the code actually consults; a pending_review queue, interrupt_before, or can_use_tool approval on high-harm paths before set_status, notify, payment, or tool execution; and per-decision records with pinned model version, inputs, outputs, and reason codes or attributions that let the result be explained or reproduced.
Example (config/oversight.yaml), before:
decision_types:
loan_denial: {auto: true}
account_freeze: {auto: true}After:
# harm = severity x probability -> human role
loan_denial: {severity: high, probability: medium, mode: human_in_the_loop}
account_freeze: {severity: high, probability: high, mode: human_in_the_loop}
fraud_flag: {severity: medium, probability: medium, mode: human_over_the_loop}
address_autofill: {severity: low, probability: low, mode: human_out_of_the_loop}Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- Singapore (SG)
- Human involvement in AI decisions should be calibrated to harm; decisions should be explainable or repeatable (Singapore MGF) Singapore Model AI Governance Framework (2nd ed.) — paras. 2.7(a), 3.13-3.15 (human involvement), 3.30 (repeatability)
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- UnitedHealth nH Predict claim-denial litigation (2023-11; alleged (not proven)). A class action filed in November 2023 alleges that UnitedHealth's nH Predict model had a 90% error rate, measured by denials reversed on appeal, while only about 0.2% of members appealed. UnitedHealth disputes the allegations; the litigation is ongoing. Source: STAT News · evidence grade: primary · cited by Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrong
- Cigna PXDX batch claim denials (reported) (2022; alleged (not proven)). ProPublica, citing internal Cigna records, reported that Cigna's PXDX system was used to reject more than 300,000 claims over two months in 2022, with physicians spending an average of 1.2 seconds on each. Cigna disputes the reporting; related lawsuits are ongoing. Source: ProPublica / The Capitol Forum · evidence grade: press of record · cited by Make human review of adverse AI decisions substantive, not nominal