Binding law — in force AI-adjacent law
Base Medicare Advantage medical necessity determinations on the enrollee's own history, physician recommendations and notes (42 CFR 422.101(c)(1)(i))
MA organizations must make medical necessity determinations based on all of: Medicare coverage and benefit criteria (and may not deny basic benefits on criteria not specified in 422.101(b) or (c)), whether the item or service is reasonable and necessary under section 1862(a)(1) of the Act, the enrollee's medical history (for example diagnoses, conditions, functional status), physician recommendations and clinical notes, and, where appropriate, the medical director's involvement (42 CFR 422.101(c)(1)(i)(A)-(D)). The rule never mentions AI; it is encoded as AI-adjacent (owner decision D-19, D-11 pattern 1) because algorithmic and AI review tools that predict outcomes from group data are where determinations without the individual record arise. Detect a utilization-review model or scoring call whose inputs never include the enrollee's own clinical record.
Trust and provenance not reviewed by a lawyer · audit-grade · source verified 4 Oct 2026 · release 2026.10.05
- Lane
- Binding law — in force In force: applies since 1 Jan 2024
- Official source
- 42 CFR 422.101(c)(1)(i) · captured 4 Oct 2026 · anchor hash (SHA-256)
19716f0548d6…· 3 more anchors in the data release - Verification
- Quoted text found word for word in the captured official document (4 Oct 2026). Source last verified 4 Oct 2026: checked against the captured official document; not in the weekly watcher's list; checked against the captured document.
- Data release
- Data release 2026.10.05, data as of 4 Oct 2026, schema 0.3.10.
- Legal review
- Not reviewed by a lawyer. TwinEthos derived this rule from the official text it cites: treat it as research to check against that text; it is not legal advice. No TwinEthos rule has been legally reviewed yet. Open questions for counsel on this rule: 1.
- Audit standard
- Audit-grade: meets all 10 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
1 detector (code pattern), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Prompt builders imported from another module
- Feature stores whose columns are not named in the code
- The clinical record may be assembled in a helper module and passed in under a generic name; trace the prompt or feature builder before reporting.
Who it applies to
- Duty falls on: insurer
- Sectors: insurance, healthcare
- Medicare Advantage organizations (42 CFR 422.2) making medical necessity determinations for MA enrollees in the United States, whatever tool, AI, algorithm or otherwise, supports the determination. The revised 422.101(c) is in force from 2023-06-05 and applies to coverage beginning 2024-01-01.
- Whether it applies depends on facts outside the code; a person has to decide.
The guard to add
Build each automated medical-necessity determination from the enrollee's own clinical record and the provider's submission, and refuse to decide on group statistics alone.
In the prompt builder or feature pipeline for each medical-necessity or coverage determination, load the enrollee's clinical history and the requesting provider's clinical documentation for this request (the attached notes, the FHIR Condition, Observation and DocumentReference resources for the member, the provider's letter of medical necessity) and pass the fields the decision needs; group or population data (a diagnosis code's typical length of stay, a cohort's approval rate, a regional benchmark) may be context but never the only input. A guard before the model or rules call raises an error, or routes the case to clinical review, when the individual clinical inputs are empty, and the inputs used are saved with the result so a reviewer or regulator can see what the determination rested on.
Where it goes: 1 application source code, 2 data models, 7 prompt construction, 9 AI output handling.
What this provision adds:
- Medicare coverage and benefit criteria and the reasonable-and-necessary test also bind each determination; basic benefits may not be denied on criteria not specified in 422.101(b) or (c).
Example (Python + Anthropic SDK), before:
features = {'cpt': req.cpt, 'icd10': req.icd10, 'cohort_approval_rate': cohort_stats(req.icd10)}
reply = client.messages.create(model=MODEL, max_tokens=400,
messages=[{'role': 'user', 'content': f'Prior authorization request: {features}'}])After:
record = member_clinical_record(req.member_id, fields=PA_CLINICAL_FIELDS) # history, conditions, meds
submission = provider_clinical_submission(req.id) # notes, letter of medical necessity
if not record or not submission:
return review_queue.enqueue(req.id, reason='individual clinical information missing')
reply = client.messages.create(model=MODEL, max_tokens=400, messages=[{'role': 'user', 'content':
render('pa_review.txt', request=req, clinical_history=record, provider_submission=submission)}])
determinations.save(req.id, inputs={'clinical_history': record.ids, 'submission': submission.ids})Control: AI coverage or medical-necessity determination not based on the individual's own clinical information. The same guard addresses 5 items with binding law in 5 jurisdictions. Engineering guidance, not legal advice.
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Meta's automated moderation over-enforced Arabic and under-enforced Hebrew content (BSR due diligence) (2021-05; disclosed by the operator). An independent human rights due diligence by BSR, commissioned and published by Meta on September 22, 2022, found that during the May 2021 Israel-Palestine escalation Arabic content saw greater over-enforcement per user than Hebrew content and Hebrew content greater under-enforcement. BSR attributes this in part to Meta having an Arabic hostile-speech classifier but no Hebrew one, and to Arabic classifiers likely being less accurate for Palestinian Arabic. BSR found no intentional bias but 'various instances of unintentional bias' with different impacts on Palestinian and Arabic-speaking users. Meta committed to implement 10 of BSR's 21 recommendations and said it had since launched a Hebrew hostile-speech classifier. Source: Meta (operator response, 2022-09-22) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits, and screen every served language
- Toxicity classifiers rate African American English as more offensive (2019; confirmed). University of Washington researchers reported at ACL 2019 that tweets in African American English and tweets by self-identified African Americans were up to two times more likely to be labelled offensive by hate-speech models trained on widely used datasets, and that Jigsaw's public Perspective API showed similar racial bias, rating AAE phrases as more toxic than non-AAE equivalents. Source: Sap et al., 'The Risk of Racial Bias in Hate Speech Detection', ACL 2019 (original researchers) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits, and screen every served language
- Ride-hailing fares higher in Chicago neighborhoods with more non-white residents (researcher audit) (2018-11; alleged (not proven)). George Washington University researchers analysing Chicago's public data on more than 100 million ride-hailing trips from November 2018 to September 2019 report that trips in neighborhoods with larger non-white populations, higher poverty, younger residents and more college-educated residents were significantly associated with higher fares. They attribute this to pricing algorithms learning from demand, supply and trip duration. The finding is an observational association from census-tract data; the pricing models themselves were not examined. Source: Pandey & Caliskan, 'Disparate Impact of Artificial Intelligence Bias in Ridehailing Economy's Price Discrimination Algorithms', AIES 2021 (original researchers) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits, and screen every served language
- Google ads suggesting arrest records served more often for Black-identifying names (2012; confirmed). Harvard researcher Latanya Sweeney searched 2,184 racially associated full names on google.com and reuters.com (a Google AdSense host) from September 24 to October 23, 2012 and found ads suggestive of an arrest record appeared more often for Black-identifying first names; on reuters.com a Black-identifying name was 25% more likely to get such an ad (statistically significant). Ads appeared regardless of whether the name had an arrest record in the advertiser's database. The paper does not determine whether the advertiser's templates or Google's click-based ad optimization caused the pattern; the advertiser, Instant Checkmate, told the author it gave Google the same ad text for groups of last names. Source: Sweeney, 'Discrimination in Online Ad Delivery' (original researcher, 2013-01-28) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits, and screen every served language
Rule id us-cms-ma-medical-necessity.determination-on-individual-medical-history · review status: primary source derived