TwinEthosRequest access

Control

AI decision system without regular accuracy/bias validation

Models used for consequential AI decisions should be regularly reviewed and validated for accuracy, relevance, and unintentional bias.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI is used without bias, fairness, or proxy-discrimination controls · control id cond.ai-decision-no-bias-validation

Reach

4items this one guard addresses
0jurisdictions where binding law on it is in force
1more where it is enacted, not yet applying
3standards and frameworks on the same control

enacted, not yet applying in European Union (EU); next date 2027-12-02.

The guard to add

Compute per-group accuracy and bias metrics in the training pipeline of every consequential-decision model, keep the results per version, and rerun them on a schedule.

A validation stage that runs whenever a decision model is trained, fine-tuned, or retrained (the same module or pipeline step as fit(), Trainer, xgb.train, or fine_tuning.jobs.create) and again on a recurring schedule against recent decisions: accuracy and error rates per group, selection rates and disparity metrics (fairlearn MetricFrame, demographic_parity_difference, AIF360 disparate impact), and drift. Results go to a versioned validation report alongside a datasheet or data card for the training data, and a threshold gate blocks promotion of a model version whose metrics regress until a named owner reviews and records a decision. The validation cadence and owner are written in the model's validation record.

Where it goes: 1 application source code, 11 CI/CD pipeline, 12 repository artifacts, 13 tests and evals.

What reviewers look for: bias or subgroup metrics computed in the training or evaluation module (fairlearn, MetricFrame, selection_rate, aif360, disparate_impact), a datasheet or data card for the training set, stored validation reports per model version, and a scheduled job or CI stage that reruns the checks and can block a release.

Example (scikit-learn + fairlearn), before:

clf = LogisticRegression(max_iter=1000).fit(X_train, y_train)
joblib.dump(clf, 'models/credit_v4.joblib')

After:

from fairlearn.metrics import MetricFrame, selection_rate, demographic_parity_difference
from sklearn.metrics import accuracy_score

clf = LogisticRegression(max_iter=1000).fit(X_train, y_train)
y_pred = clf.predict(X_test)
mf = MetricFrame(metrics={'accuracy': accuracy_score, 'selection_rate': selection_rate},
                 y_true=y_test, y_pred=y_pred, sensitive_features=A_test)
dpd = demographic_parity_difference(y_test, y_pred, sensitive_features=A_test)
write_validation_report('credit_v4', mf.by_group, dpd)
if dpd > MAX_DPD:
    raise SystemExit('bias gate failed: owner review required before release')
joblib.dump(clf, 'models/credit_v4.joblib')

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Upcoming dates

Every rule this guard addresses

Binding law — not yet in force or stayed (1)

Standard / soft law (3)

Related incidents

No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.