Control
GenAI without published training-data summary
A generative AI developer must publish a high-level summary of training datasets.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
Law in force in California (US-CA).
The guard to add
Publish a high-level training-data summary for each generative AI system on the developer's website, and revise it whenever the system is retrained or substantially modified.
A training-data disclosure document the developer publishes on its website and keeps in the repository (for example docs/training-data-disclosure.md or the training-data section of a model card or datasheet), describing at a high level each dataset used to train the generative system: where it came from, what it contains, how it was collected and processed, and its rights and personal-data status. The model or data team owns it and revises it whenever the system is released, retrained, or substantially modified (new datasets, fine-tuning, synthetic data). Where the repository holds a dataset manifest, a CI check fails a change to the manifest that does not also update the disclosure.
Where it goes: 12 repository artifacts, 14 user-facing text, 11 CI/CD pipeline.
What reviewers look for: a published training-data summary reachable from the developer's site that matches the current dataset manifest; a dated revision for each release or substantial modification; and, where datasets are tracked in the repo, a check tying manifest changes to disclosure updates rather than a one-off document.
Organizational control: the evidence is a kept record, its owner and its upkeep, not code.
Example (docs/training-data-disclosure.md), before:
# Model
Trained on a large corpus of public and licensed data.After:
# Training data summary (v3, updated 2026-09-01)
## web-crawl-2025
Source/owner: own crawler, public pages. Purpose: general language ability.
Size: ~1.2B documents. Types: text. Licensed: no. Personal information: may be present; PII filtered.
Collected: 2023-01 to 2025-06. First used: 2025-08. Synthetic: no.
## support-dialogs-licensed
Source/owner: Acme Corp, licensed. Size: 4M dialogs. Synthetic: 20% paraphrased.Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Binding law — in force (1)
- California (US-CA)
- GenAI developers must publish a training-data summary Cal. Civ. Code 3111(a)