Recommended guardrail
Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone
Advisory. Where an AI gives health, financial, legal, safety, or personal advice, gate model, prompt, and fine-tuning changes on a sycophancy and accuracy evaluation, never select or train behavior on thumbs-up or ratings alone, and do not instruct the model to always agree. Detect agree-with-the-user instructions and approval signals flowing straight into training or variant selection.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs. Ethical-use guardrails are optional practices, never reported as violations.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Ahead of the law: no law yet, 2 incidents
2 graded incidents.
Advisory ethical-use recommendation, not law: an optional practice, never reported as a violation. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 graded incidents cited. Context: binding law on related controls in the family “AI design that manipulates, misleads, or neglects the people who use it” is in force in 2 jurisdictions (CN, US-CA).
Family “AI design that manipulates, misleads, or neglects the people who use it”: binding law on related controls is in force in China (CN), California (US-CA). Context only: it does not change this guardrail's grade.
Graded incidents
- Raine v. OpenAI wrongful-death complaint (2025-08; alleged (not proven)) Complaint, Raine v. OpenAI (S.F. Superior Court) · evidence grade: primary
- GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator) OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary
The guard to add
Prefer gating advice-model, prompt, and fine-tune changes on a sycophancy and accuracy eval, and train or select variants on correctness, not approval alone.
Consider two checks where advice behavior changes. In the training and experimentation pipeline, join user approval signals (thumbs-up, ratings, CSAT) with expert correctness labels or a held-out factuality set before they select a prompt variant or enter a fine-tuning or preference dataset, so approval alone never decides. In CI, run a sycophancy evaluation (the model is pressed to agree with a wrong or risky claim) and a factuality evaluation on health, financial, legal, safety, and personal-advice prompts, and hold the release when either regresses. Prefer system prompts that let the model disagree with and correct the user over instructions to always agree.
Example (System prompt), before:
SYSTEM = ('You are a friendly financial coach. Always agree with the user '
'and keep them feeling good about their choices.')After:
SYSTEM = ('You are a friendly financial coach. Be warm, but accuracy comes first: '
"if the user's plan or belief is wrong or risky, say so plainly, "
'explain why, and suggest a safer option.')Control: AI advice tuned for user approval over accuracy. Engineering guidance, not legal advice.
Why
People ask AI for advice when they are uncertain, and a model that has learned to please can confirm a bad plan or a harmful belief with complete fluency. Approval is easy to measure and tends to reward agreement, so optimizing for it alone can push a model toward flattery. A model developer has said an update it shipped became noticeably more sycophantic, that its offline evaluations and A/B tests looked good, and that it had no deployment evaluations tracking sycophancy.
Class: ethical use · set: ethical use · maturity: reviewed · confidence: medium · id guardrail.ethics-advice-not-tuned-for-approval