Standard / soft law
Health AI/LMMs should keep a clinician accountable with oversight and redress (WHO)
Per the WHO Guidance on AI for Health (LMMs), under the consensus principles 'foster responsibility and accountability' and 'protect autonomy', AI/LMMs used in health care must be subject to human oversight — a health-care provider or clinician remains accountable for clinical decisions and must be able to review, override, and take responsibility for AI outputs — with clear assignment of responsibility and mechanisms for redress for individuals harmed by an AI-informed decision. Detect a health-AI decision path with no clinician oversight/override or no accountability/redress mechanism.
Who it applies to
- Duty falls on: developer, deployer
- Systems covered: automated decision, high risk
- Sectors: healthcare
- Developers, providers, and deployers of AI/LMMs used for health care, medical, or public-health purposes. Voluntary WHO guidance; governments have primary responsibility to set standards. Globally influential health-AI baseline.
- Whether it applies depends on facts outside the code; a person has to decide.
The guard to add
Hold AI-generated clinical output as a draft until an accountable clinician reviews and signs it, and record who approved it before it reaches the chart or the patient.
A clinician sign-off step between the model call and every clinical sink: AI-drafted notes, summaries, diagnostic suggestions, triage levels, and treatment plans are stored as drafts (FHIR DocumentReference.docStatus 'preliminary', DiagnosticReport.status 'preliminary', CarePlan.status 'draft') and become final, active, or visible to the patient only through an action by an authorized clinician that records reviewed_by and reviewed_at. Configuration flags that auto-sign or auto-finalize AI-drafted records stay false, and provenance shows the AI as a contributing device and the clinician as verifier. The deployment also names who is accountable for AI-assisted decisions and gives patients a complaint or redress route.
Where it goes: 1 application source code, 2 data models, 9 AI output handling, 3 config and feature flags.
Example (Python + OpenAI SDK + FHIR REST), before:
note = client.chat.completions.create(model=MODEL, messages=msgs).choices[0].message.content
requests.post(f'{FHIR_BASE}/DocumentReference', json=doc_ref(patient_id, note, doc_status='final'))After:
note = client.chat.completions.create(model=MODEL, messages=msgs).choices[0].message.content
requests.post(f'{FHIR_BASE}/DocumentReference',
json=doc_ref(patient_id, note, doc_status='preliminary')) # AI draft
def practitioner_review_and_sign(doc_id, practitioner): # only path to 'final'
doc = requests.get(f'{FHIR_BASE}/DocumentReference/{doc_id}').json()
doc['docStatus'] = 'final'
doc['authenticator'] = {'reference': f'Practitioner/{practitioner.id}'}
requests.put(f'{FHIR_BASE}/DocumentReference/{doc_id}', json=doc)
audit.record(doc_id, reviewed_by=practitioner.id, reviewed_at=utcnow())Control: Health AI without clinician oversight/accountability + redress. The same guard addresses 3 items with binding law in 2 jurisdictions. Engineering guidance, not legal advice.
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- UnitedHealth nH Predict claim-denial litigation (2023-11; alleged (not proven)). A class action filed in November 2023 alleges that UnitedHealth's nH Predict model had a 90% error rate, measured by denials reversed on appeal, while only about 0.2% of members appealed. UnitedHealth disputes the allegations; the litigation is ongoing. Source: STAT News · evidence grade: primary · cited by Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrong
- Cigna PXDX batch claim denials (reported) (2022; alleged (not proven)). ProPublica, citing internal Cigna records, reported that Cigna's PXDX system was used to reject more than 300,000 claims over two months in 2022, with physicians spending an average of 1.2 seconds on each. Cigna disputes the reporting; related lawsuits are ongoing. Source: ProPublica / The Capitol Forum · evidence grade: press of record · cited by Make human review of adverse AI decisions substantive, not nominal
Rule id who-health-ai.clinician-oversight-accountability · review status: primary source derived