Control
GenAI training data without lawful source / IP / consent controls
GenAI providers must use lawfully-sourced data and foundational models, not infringe IP, obtain consent for personal information, and ensure training-data quality.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
Law in force in China (CN).
The guard to add
Gate every dataset entering pre-training or fine-tuning on a recorded lawful source and licence, and consent-filter or scrub personal information first.
A gate in the data pipeline between collection (load_dataset, crawlers, Common Crawl WARC readers, exports of user chats) and training (Trainer, SFTTrainer, fine_tuning.jobs.create) that keeps only records whose source and licence are on an allowlist, checks robots.txt before crawling, drops user content without a training-consent flag, and runs PII detection and anonymisation on the rest. It writes a provenance manifest per shard (source, licence, retrieval date, consent basis, filters applied) kept with the model version, alongside the IP clearance record and the data-quality measures applied. Base models you fine-tune get the same source and licence entry.
Where it goes: 1 application source code, 11 CI/CD pipeline, 12 repository artifacts.
What reviewers look for: on every path from a dataset loader, crawler or chat-log export to a trainer or fine-tuning job, a licence allowlist filter, a robots.txt check for crawls, a consented_to_training filter or Presidio-style scrub for personal information, and a provenance manifest written per shard and stored with the model version.
Example (Hugging Face datasets + TRL + Presidio), before:
ds = load_dataset('json', data_files='crawl/*.jsonl', split='train')
trainer = SFTTrainer(model=model, train_dataset=ds)
trainer.train()After:
ALLOWED_LICENSES = {'cc-by-4.0', 'cc0-1.0', 'apache-2.0', 'licensed-by-contract'}
analyzer, anonymizer = AnalyzerEngine(), AnonymizerEngine()
def scrub(row):
hits = analyzer.analyze(text=row['text'], language=LANG)
row['text'] = anonymizer.anonymize(text=row['text'], analyzer_results=hits).text
return row
ds = load_dataset('json', data_files='crawl/*.jsonl', split='train')
ds = ds.filter(lambda r: r['license'] in ALLOWED_LICENSES).map(scrub)
write_provenance_manifest(ds, out='manifests/shard-000.json')
trainer = SFTTrainer(model=model, train_dataset=ds)
trainer.train()Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Binding law — in force (1)
- China (CN)
- GenAI training data must have lawful sources, cleared IP, and consent (China) 生成式人工智能服务管理暂行办法 第七条 (Art. 7)