Binding law — in force
GenAI developers must publish a training-data summary
Developers of generative AI made publicly available to Californians (released/substantially-modified on or after 2022-01-01) must post on their website a high-level summary of the training datasets covering 12 enumerated items: sources/owners, purpose, data-point counts, data types, copyright/trademark/patent status, purchased/licensed status, whether personal information or aggregate consumer information is included, cleaning/processing, collection time period, first-use dates, and synthetic-data use. Detect a generative AI codebase with no training-data-disclosure artifact.
Who it applies to
- Duty falls on: developer
- Developers (incl. those who substantially modify) of generative AI made available to Californians. Applies to systems released/substantially-modified on or after 2022-01-01. Exemptions: security/integrity-only, national-airspace operation, or federal-only national-security/defense systems.
The guard to add
Organizational artifact to keep (not verifiable from code); the guard is the record, its owner and its upkeep.
Publish a high-level training-data summary for each generative AI system on the developer's website, and revise it whenever the system is retrained or substantially modified.
A training-data disclosure document the developer publishes on its website and keeps in the repository (for example docs/training-data-disclosure.md or the training-data section of a model card or datasheet), describing at a high level each dataset used to train the generative system: where it came from, what it contains, how it was collected and processed, and its rights and personal-data status. The model or data team owns it and revises it whenever the system is released, retrained, or substantially modified (new datasets, fine-tuning, synthetic data). Where the repository holds a dataset manifest, a CI check fails a change to the manifest that does not also update the disclosure.
Where it goes: 12 repository artifacts, 14 user-facing text, 11 CI/CD pipeline.
What this provision adds:
- Cover the enumerated items: sources/owners, purpose, data-point counts, data types, copyright/trademark/patent status, purchased/licensed status, personal or aggregate consumer information, cleaning/processing, collection period, first-use dates, synthetic-data use.
- Post the summary on the developer's website for each generative system made available to Californians and released or substantially modified on or after 2022-01-01.
Example (docs/training-data-disclosure.md), before:
# Model
Trained on a large corpus of public and licensed data.After:
# Training data summary (v3, updated 2026-09-01)
## web-crawl-2025
Source/owner: own crawler, public pages. Purpose: general language ability.
Size: ~1.2B documents. Types: text. Licensed: no. Personal information: may be present; PII filtered.
Collected: 2023-01 to 2025-06. First used: 2025-08. Synthetic: no.
## support-dialogs-licensed
Source/owner: Acme Corp, licensed. Size: 4M dialogs. Synthetic: 20% paraphrased.Control: GenAI without published training-data summary. The same guard addresses 1 item with binding law in 1 jurisdiction. Engineering guidance, not legal advice.
Rule id ca-ab2013.training-data-disclosure · review status: primary source derived