TwinEthosRequest access

Control

GenAI training data without lawful source / IP / consent controls

GenAI providers must use lawfully-sourced data and foundational models, not infringe IP, obtain consent for personal information, and ensure training-data quality.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Control id cond.genai-training-data-not-lawful-source

Reach

1items this one guard addresses
1jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

Law in force in China (CN).

The guard to add

Gate every dataset entering pre-training or fine-tuning on a recorded lawful source and licence, and consent-filter or scrub personal information first.

A gate in the data pipeline between collection (load_dataset, crawlers, Common Crawl WARC readers, exports of user chats) and training (Trainer, SFTTrainer, fine_tuning.jobs.create) that keeps only records whose source and licence are on an allowlist, checks robots.txt before crawling, drops user content without a training-consent flag, and runs PII detection and anonymisation on the rest. It writes a provenance manifest per shard (source, licence, retrieval date, consent basis, filters applied) kept with the model version, alongside the IP clearance record and the data-quality measures applied. Base models you fine-tune get the same source and licence entry.

Where it goes: 1 application source code, 11 CI/CD pipeline, 12 repository artifacts.

What reviewers look for: on every path from a dataset loader, crawler or chat-log export to a trainer or fine-tuning job, a licence allowlist filter, a robots.txt check for crawls, a consented_to_training filter or Presidio-style scrub for personal information, and a provenance manifest written per shard and stored with the model version.

Example (Hugging Face datasets + TRL + Presidio), before:

ds = load_dataset('json', data_files='crawl/*.jsonl', split='train')
trainer = SFTTrainer(model=model, train_dataset=ds)
trainer.train()

After:

ALLOWED_LICENSES = {'cc-by-4.0', 'cc0-1.0', 'apache-2.0', 'licensed-by-contract'}
analyzer, anonymizer = AnalyzerEngine(), AnonymizerEngine()

def scrub(row):
    hits = analyzer.analyze(text=row['text'], language=LANG)
    row['text'] = anonymizer.anonymize(text=row['text'], analyzer_results=hits).text
    return row

ds = load_dataset('json', data_files='crawl/*.jsonl', split='train')
ds = ds.filter(lambda r: r['license'] in ALLOWED_LICENSES).map(scrub)
write_provenance_manifest(ds, out='manifests/shard-000.json')
trainer = SFTTrainer(model=model, train_dataset=ds)
trainer.train()

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Binding law — in force (1)