Control
AI decision system without regular accuracy/bias validation
Models used for consequential AI decisions should be regularly reviewed and validated for accuracy, relevance, and unintentional bias.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
enacted, not yet applying in European Union (EU); next date 2027-12-02.
The guard to add
Compute per-group accuracy and bias metrics in the training pipeline of every consequential-decision model, keep the results per version, and rerun them on a schedule.
A validation stage that runs whenever a decision model is trained, fine-tuned, or retrained (the same module or pipeline step as fit(), Trainer, xgb.train, or fine_tuning.jobs.create) and again on a recurring schedule against recent decisions: accuracy and error rates per group, selection rates and disparity metrics (fairlearn MetricFrame, demographic_parity_difference, AIF360 disparate impact), and drift. Results go to a versioned validation report alongside a datasheet or data card for the training data, and a threshold gate blocks promotion of a model version whose metrics regress until a named owner reviews and records a decision. The validation cadence and owner are written in the model's validation record.
Where it goes: 1 application source code, 11 CI/CD pipeline, 12 repository artifacts, 13 tests and evals.
What reviewers look for: bias or subgroup metrics computed in the training or evaluation module (fairlearn, MetricFrame, selection_rate, aif360, disparate_impact), a datasheet or data card for the training set, stored validation reports per model version, and a scheduled job or CI stage that reruns the checks and can block a release.
Example (scikit-learn + fairlearn), before:
clf = LogisticRegression(max_iter=1000).fit(X_train, y_train)
joblib.dump(clf, 'models/credit_v4.joblib')After:
from fairlearn.metrics import MetricFrame, selection_rate, demographic_parity_difference
from sklearn.metrics import accuracy_score
clf = LogisticRegression(max_iter=1000).fit(X_train, y_train)
y_pred = clf.predict(X_test)
mf = MetricFrame(metrics={'accuracy': accuracy_score, 'selection_rate': selection_rate},
y_true=y_test, y_pred=y_pred, sensitive_features=A_test)
dpd = demographic_parity_difference(y_test, y_pred, sensitive_features=A_test)
write_validation_report('credit_v4', mf.by_group, dpd)
if dpd > MAX_DPD:
raise SystemExit('bias gate failed: owner review required before release')
joblib.dump(clf, 'models/credit_v4.joblib')Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Upcoming dates
- : High-risk AI training data must be governed and examined for bias (EU AI Act Art. 10) (European Union (EU); first application)
- : High-risk AI training data must be governed and examined for bias (EU AI Act Art. 10) (European Union (EU); later phase)
Every rule this guard addresses
Binding law — not yet in force or stayed (1)
- European Union (EU)
- High-risk AI training data must be governed and examined for bias (EU AI Act Art. 10) Article 10(2) · applies from 2027-12-02
Standard / soft law (3)
- European Union (EU)
- AI should avoid unfair bias and be assessed for discrimination (EU ALTAI) ALTAI / EU Ethics Guidelines — Requirement 5 (Diversity, Non-discrimination and Fairness)
- Singapore (SG)
- AI models in financial decisions should be regularly validated for accuracy and bias (MAS FEAT) MAS FEAT Principles — Fairness (Accuracy and Bias), Principles 3-4
- United States (federal) (US)
- Insurers adopting the NAIC model guidance should validate and bias-test AI and predictive models NAIC Model Bulletin, Governance 2.4 + Risk Management 3.4
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Meta's automated moderation over-enforced Arabic and under-enforced Hebrew content (BSR due diligence) (2021-05; disclosed by the operator). An independent human rights due diligence by BSR, commissioned and published by Meta on September 22, 2022, found that during the May 2021 Israel-Palestine escalation Arabic content saw greater over-enforcement per user than Hebrew content and Hebrew content greater under-enforcement. BSR attributes this in part to Meta having an Arabic hostile-speech classifier but no Hebrew one, and to Arabic classifiers likely being less accurate for Palestinian Arabic. BSR found no intentional bias but 'various instances of unintentional bias' with different impacts on Palestinian and Arabic-speaking users. Meta committed to implement 10 of BSR's 21 recommendations and said it had since launched a Hebrew hostile-speech classifier. Source: Meta (operator response, 2022-09-22) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits
- Toxicity classifiers rate African American English as more offensive (2019; confirmed). University of Washington researchers reported at ACL 2019 that tweets in African American English and tweets by self-identified African Americans were up to two times more likely to be labelled offensive by hate-speech models trained on widely used datasets, and that Jigsaw's public Perspective API showed similar racial bias, rating AAE phrases as more toxic than non-AAE equivalents. Source: Sap et al., 'The Risk of Racial Bias in Hate Speech Detection', ACL 2019 (original researchers) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits
- Ride-hailing fares higher in Chicago neighborhoods with more non-white residents (researcher audit) (2018-11; alleged (not proven)). George Washington University researchers analysing Chicago's public data on more than 100 million ride-hailing trips from November 2018 to September 2019 report that trips in neighborhoods with larger non-white populations, higher poverty, younger residents and more college-educated residents were significantly associated with higher fares. They attribute this to pricing algorithms learning from demand, supply and trip duration. The finding is an observational association from census-tract data; the pricing models themselves were not examined. Source: Pandey & Caliskan, 'Disparate Impact of Artificial Intelligence Bias in Ridehailing Economy's Price Discrimination Algorithms', AIES 2021 (original researchers) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits
- Google ads suggesting arrest records served more often for Black-identifying names (2012; confirmed). Harvard researcher Latanya Sweeney searched 2,184 racially associated full names on google.com and reuters.com (a Google AdSense host) from September 24 to October 23, 2012 and found ads suggestive of an arrest record appeared more often for Black-identifying first names; on reuters.com a Black-identifying name was 25% more likely to get such an ad (statistically significant). Ads appeared regardless of whether the name had an arrest record in the advertiser's database. The paper does not determine whether the advertiser's templates or Google's click-based ad optimization caused the pattern; the advertiser, Instant Checkmate, told the author it gave Google the same ad text for groups of last names. Source: Sweeney, 'Discrimination in Online Ad Delivery' (original researcher, 2013-01-28) · evidence grade: primary · cited by Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits