Binding law — in force
Algorithmic review must use physician-set criteria, verified by board-certified physicians before changes and after errors (Illinois HB 2472)
From 2025-01-01, utilization review programs that use algorithmic automated processes to decide whether to render adverse determinations based on medical necessity must use objective, evidence-based criteria compliant with URAC or NCQA accreditation requirements and prove compliance with each registration (215 ILCS 134/85(b-10), (a)); the registration must attach policies and procedures (1) ensuring that licensed physicians with relevant board certifications establish all criteria the process uses and (2) for a program integrity system that, before new or revised criteria are used and when implementation errors are found, requires such physicians to verify that the process and its corrections yield results consistent with the criteria for their certified field. A plan or program using an automated process must have that accreditation and those policies (45(i)). Detect the absence of physician criteria sign-off and pre-release integrity verification.
Trust and provenance not reviewed by a lawyer · audit-grade · source verified 3 Oct 2026 · release 2026.10.03.4
- Lane
- Binding law — in force In force: applies since 1 Jan 2025
- Official source
- 215 ILCS 134/85(b-10) · captured 3 Oct 2026 · anchor hash (SHA-256)
9be9ea68cc49…· 12 more anchors in the data release - Verification
- Quoted text found word for word in the captured official document (3 Oct 2026). Source last verified 3 Oct 2026: checked against the captured official document.
- Data release
- Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.9.
- Legal review
- Not reviewed by a lawyer. TwinEthos derived this rule from the official text it cites: treat it as research to check against that text; it is not legal advice. No TwinEthos rule has been legally reviewed yet. Open questions for counsel on this rule: 1.
- Audit standard
- Audit-grade: meets all 10 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
1 detector (missing artifact), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Verification run in a vendor's environment
Who it applies to
- Duty falls on: insurer, organization
- Sectors: insurance, healthcare
- Health care plans (HMOs, managed care community networks and accountable care entities with a provider network) and any person conducting a utilization review program in Illinois (registered with the Department of Insurance) that uses an algorithmic automated process in utilization review for medical necessity. In force 2025-01-01.
- Not covered:
- Section 85 does not apply to persons providing utilization review program services only to the federal government (215 ILCS 134/85(c)(1))
- Section 85 does not apply to self-insured ERISA health plans, though it applies to persons conducting utilization review on their behalf (85(c)(2))
- Section 85 does not apply to hospitals and medical groups performing utilization review for internal purposes unless conducted for another person (85(c)(3))
- 'Health care plan' (Section 45) excludes indemnity policies (including those with a contracted network), dental-only and vision-only plans, preferred provider administrators, self-insured ERISA plans, workers' compensation care and certain union-affiliated plans (134/10)
- Whether it applies depends on facts outside the code; a person has to decide.
The guard to add
Pin dated model versions and run a blocking behavior and safety eval in CI whenever model ids, prompts, or inference settings change, then keep monitoring quality in production.
Reference models by dated snapshot ids in one config file instead of floating aliases (-latest, -preview, undated names), so the model changes only through a commit. A CI job triggered by changes to that file, the prompt files, and inference settings (sampling, quantization, routing, provider) runs the behavior and safety eval suite (promptfoo, deepeval, inspect_ai, or openai/evals) and blocks merge or release on regression. A scheduled online eval or canary, plus quality and refusal-rate alerts on production gen_ai spans, catches provider-side changes that no pre-release gate can see.
Where it goes: 3 config and feature flags, 8 model configuration, 11 CI/CD pipeline, 13 tests and evals.
What this provision adds:
- Attach the criteria and program-integrity policies to the Department registration (every 2 years) and to each renewal.
Example (App model config), before:
llm:
model: gpt-4o
temperature: 0.7After:
llm:
model: gpt-4o-2024-08-06 # change only via PR; triggers the eval workflow
temperature: 0.7Control: Model, version, or serving change reaches users without re-running behavior and safety evaluations. The same guard addresses 3 items with binding law in 1 jurisdiction. Engineering guidance, not legal advice.
Standards that recommend the same control
- GenAI risks should be re-assessed and red-teamed after fine-tuning, RAG, or a new use of a third-party model (NIST GenAI Profile) (NIST GenAI Profile (AI 600-1) · NIST AI 600-1, MG-3.1-003)
Related incidents
- Serving-stack changes silently degraded Claude output quality (2025-08; disclosed by the operator). Anthropic reports that three infrastructure bugs, including one introduced by a runtime performance optimization, intermittently degraded Claude's responses between early August and early September 2025, and that its benchmarks, safety evaluations, and canary deployments did not capture the degradation. It now runs quality evaluations continuously on production systems. Source: Anthropic (operator postmortem, 2025-09-17) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users
- GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator). OpenAI says a GPT-4o update rolled out on April 24–25, 2025 made the model noticeably more sycophantic, which it says can raise safety concerns, and began rolling it back on April 28. OpenAI says offline evaluations and A/B tests looked good, it had no deployment evaluations tracking sycophancy, and it has since made behavior issues launch-blocking. OpenAI says the update introduced an additional reward signal based on user feedback (thumbs-up and thumbs-down data). Source: OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users
Rule id il-hb2472.automated-process-criteria-and-integrity · review status: primary source derived