Recommended guardrail
Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits
Advisory. When personal names, postal codes or census tracts, language or dialect, or device or IP geolocation are features in AI ranking, pricing, content moderation, or ad targeting, compare outcomes across the groups those features can stand in for before launch and on a schedule, and mitigate or document any disparity. Detect proxy-capable features flowing into those models with no disparity evaluation on the path, and feature lists that name them.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs. Ethical-use guardrails are optional practices, never reported as violations.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law in force in 2 jurisdictions
Law in force in 2 jurisdictions · 3 standards and frameworks · 4 graded incidents.
Advisory ethical-use recommendation, not law: an optional practice, never reported as a violation. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is in force in 2 jurisdictions (US-CO, US-IL). 3 standards and frameworks recommend it (IEEE 7003-2024, MAS FEAT Principles, NIST AI RMF 1.0). 4 graded incidents cited.
Law in force on this control or cited as convergence
- Illinois employers must not use AI that discriminates on protected classes or uses ZIP codes as a proxy (employment decisions) (Illinois (US-IL); 775 ILCS 5/2-102(L)(1); cited)
- Insurance algorithms must not unfairly discriminate via external data (Colorado (US-CO); 3 CCR 702-10, Regulation 10-1-1, Section 5.A; cited)
Standards and frameworks
- Systems should document a proxy-variable / algorithmic-bias analysis (IEEE 7003-2024; IEEE 7003-2024 Clause 5 (bias profiling); cited)
- AI decision systems should have a documented fairness/bias evaluation (United States (federal) (US); NIST AI RMF 1.0 — MEASURE 2.11; cited)
- Personal-attribute inputs to AI financial decisions should be justified; no unjustified systematic disadvantage (MAS FEAT) (Singapore (SG); MAS FEAT Principles — Fairness (Justifiability), Principles 1-2; cited)
Family “AI is used without bias, fairness, or proxy-discrimination controls”: binding law on related controls is in force in Colorado (US-CO), Illinois (US-IL), New York City (US-NY-NYC), Texas (US-TX); enacted, not yet applying in European Union (EU). Context only: it does not change this guardrail's grade.
Graded incidents
- Meta's automated moderation over-enforced Arabic and under-enforced Hebrew content (BSR due diligence) (2021-05; disclosed by the operator) Meta (operator response, 2022-09-22) · evidence grade: primary
- Toxicity classifiers rate African American English as more offensive (2019; confirmed) Sap et al., 'The Risk of Racial Bias in Hate Speech Detection', ACL 2019 (original researchers) · evidence grade: primary
- Ride-hailing fares higher in Chicago neighborhoods with more non-white residents (researcher audit) (2018-11; alleged (not proven)) Pandey & Caliskan, 'Disparate Impact of Artificial Intelligence Bias in Ridehailing Economy's Price Discrimination Algorithms', AIES 2021 (original researchers) · evidence grade: primary
- Google ads suggesting arrest records served more often for Black-identifying names (2012; confirmed) Sweeney, 'Discrimination in Online Ad Delivery' (original researcher, 2013-01-28) · evidence grade: primary
The guard to add
Consider comparing outcomes across the groups that name, postal-code, language, or geolocation features can stand in for, before launch and on a schedule.
Consider adding a disparity evaluation wherever proxy-capable features (personal names, ZIP or postal code or census tract, language, locale or dialect, IP or device geolocation) feed a model that prices, ranks, moderates, or targets ads: in the training or evaluation pipeline and as a scheduled job on live outcomes, compare price, rank position, removal rate, or ad delivery across the proxied groups (for example via census-tract demographics or BISG-inferred race and ethnicity, using fairlearn MetricFrame with sensitive_features). Prefer recording each proxy feature in the model card or a feature-review record with the measured disparity and the outcome: mitigated, coarsened, removed, or justified. Removing or coarsening the feature after a documented review is a reasonable alternative to ongoing measurement.
Example (scikit-learn pricing model + fairlearn), before:
features = ['basket_size', 'zip_code', 'device_type', 'visits_30d']
model = GradientBoostingRegressor().fit(train[features], train['price_multiplier'])After:
features = ['basket_size', 'zip_code', 'device_type', 'visits_30d']
model = GradientBoostingRegressor().fit(train[features], train['price_multiplier'])
# compare quoted prices across groups the ZIP feature can stand in for
groups = tract_demographics.majority_group(test['census_tract'])
mf = MetricFrame(metrics={'mean_price': lambda y_t, y_p: y_p.mean()},
y_true=test['price_multiplier'], y_pred=model.predict(test[features]),
sensitive_features=groups)
record_feature_review('zip_code', by_group=mf.by_group, gap=mf.difference())Control: Protected-attribute proxies used as AI features without a disparity evaluation. Engineering guidance, not legal advice.
Why
A postal code, a name, or the language someone writes in carries information about race, ethnicity, and national origin whether or not anyone intends it, and a model optimizing price, reach, or enforcement will use that information if it helps the objective. Outside the sectors where anti-discrimination law reaches automated decisions, few rules ask anyone to look, so disparities can persist unnoticed. A responsible integration measures outcomes by group wherever a proxy-capable feature is in play. Researchers have reported that ads suggesting an arrest record appeared more often for Black-identifying names, and that hate-speech classifiers trained on widely used datasets, and a widely used toxicity API, rated African American English as more offensive; an independent assessment commissioned by Meta found more over-enforcement of Arabic than Hebrew content during a May 2021 escalation, in part because its classifiers differed by language; and researchers analysing public trip data report that ride-hailing fares were higher in Chicago neighborhoods with more non-white residents, an association that did not examine the pricing models themselves.
Class: ethical use · set: ethical use · maturity: reviewed · confidence: medium · id guardrail.ethics-evaluate-proxy-features-for-disparity