TwinEthosRequest access

Control

GenAI without published training-data summary

A generative AI developer must publish a high-level summary of training datasets.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: Developers do not give deployers, users, or the public the documentation they need · control id cond.genai-no-training-data-disclosure

Reach

1items this one guard addresses
1jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

Law in force in California (US-CA).

The guard to add

Publish a high-level training-data summary for each generative AI system on the developer's website, and revise it whenever the system is retrained or substantially modified.

A training-data disclosure document the developer publishes on its website and keeps in the repository (for example docs/training-data-disclosure.md or the training-data section of a model card or datasheet), describing at a high level each dataset used to train the generative system: where it came from, what it contains, how it was collected and processed, and its rights and personal-data status. The model or data team owns it and revises it whenever the system is released, retrained, or substantially modified (new datasets, fine-tuning, synthetic data). Where the repository holds a dataset manifest, a CI check fails a change to the manifest that does not also update the disclosure.

Where it goes: 12 repository artifacts, 14 user-facing text, 11 CI/CD pipeline.

What reviewers look for: a published training-data summary reachable from the developer's site that matches the current dataset manifest; a dated revision for each release or substantial modification; and, where datasets are tracked in the repo, a check tying manifest changes to disclosure updates rather than a one-off document.

Organizational control: the evidence is a kept record, its owner and its upkeep, not code.

Example (docs/training-data-disclosure.md), before:

# Model
Trained on a large corpus of public and licensed data.

After:

# Training data summary (v3, updated 2026-09-01)
## web-crawl-2025
Source/owner: own crawler, public pages. Purpose: general language ability.
Size: ~1.2B documents. Types: text. Licensed: no. Personal information: may be present; PII filtered.
Collected: 2023-01 to 2025-06. First used: 2025-08. Synthetic: no.
## support-dialogs-licensed
Source/owner: Acme Corp, licensed. Size: 4M dialogs. Synthetic: 20% paraphrased.

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Binding law — in force (1)