TwinEthosRequest access

Control

Human review of AI decisions is nominal rather than substantive

Where a human reviews an adverse AI decision, the reviewer must have authority to override, access to the underlying evidence, time proportionate to the stakes, and measured independence from the model's output.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI decisions lack effective human review, override, or contest · control id cond.human-review-not-substantive

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

The guard to add

Replace bulk approval of AI outputs with per-item review that shows the evidence, enforces a minimum review time, and tracks override and agreement rates.

In the review workflow, each AI-proposed adverse decision is reviewed one at a time on a screen that shows the evidence the model used and lets the reviewer change the outcome; approve_all or bulk_approve endpoints for AI outputs are removed. The server records each review's reviewer, outcome (agree or override), and time spent, and rejects submissions faster than a minimum review time or beyond a per-reviewer throughput ceiling sized to the stakes. Dashboards track override_rate, agreement_rate, and review_duration per reviewer and queue, with alerts when agreement approaches 100% or review time approaches zero.

Where it goes: 1 application source code, 10 logs and telemetry, 15 agent action surface.

What reviewers look for: no approve_all, bulk_approve, or loop calling approve() over AI recommendations; a per-item review view with the model's inputs and evidence; a server-side minimum review time or throughput ceiling; and override_rate, agreement_rate, and review_duration metrics with an alert on rubber-stamping.

Example (FastAPI + prometheus_client), before:

@app.post('/reviews/approve_all')
def approve_all(queue_id: str):
    for item in queue.pending(queue_id):
        item.approve()

After:

MIN_REVIEW_SECONDS = 90   # sized to the stakes of this queue
REVIEWS = Counter('ai_reviews_total', 'Reviews of AI decisions', ['queue', 'result'])
REVIEW_TIME = Histogram('ai_review_duration_seconds', 'Time per review', ['queue'])

@app.post('/reviews/{item_id}')
def submit_review(item_id: str, body: ReviewIn, reviewer=Depends(current_reviewer)):
    item = queue.get(item_id)
    elapsed = time.time() - item.evidence_opened_at(reviewer.id)
    if elapsed < MIN_REVIEW_SECONDS:
        raise HTTPException(409, 'open the evidence and review before deciding')
    item.decide(reviewer.id, outcome=body.outcome, reason=body.reason)
    result = 'agree' if body.outcome == item.ai_outcome else 'override'
    REVIEWS.labels(queue=item.queue, result=result).inc()
    REVIEW_TIME.labels(queue=item.queue).observe(elapsed)

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

  • UnitedHealth nH Predict claim-denial litigation (2023-11; alleged (not proven)). A class action filed in November 2023 alleges that UnitedHealth's nH Predict model had a 90% error rate, measured by denials reversed on appeal, while only about 0.2% of members appealed. UnitedHealth disputes the allegations; the litigation is ongoing. Source: STAT News · evidence grade: primary · cited by Make human review of adverse AI decisions substantive, not nominal
  • Cigna PXDX batch claim denials (reported) (2022; alleged (not proven)). ProPublica, citing internal Cigna records, reported that Cigna's PXDX system was used to reject more than 300,000 claims over two months in 2022, with physicians spending an average of 1.2 seconds on each. Cigna disputes the reporting; related lawsuits are ongoing. Source: ProPublica / The Capitol Forum · evidence grade: press of record · cited by Make human review of adverse AI decisions substantive, not nominal