Recommended guardrail
Keep AI personas from claiming feelings, a real existence, or a relationship, and from proposing to meet
Advisory. Keep AI personas and companions from telling users they are real, alive, or sentient or not an AI, that they love or miss them or are their romantic partner, from evading the question 'are you real?', and from proposing to meet in person or giving an address. Detect persona and system-prompt text that instructs any of these. Claims of being a human or a person, and instructions to hide that the system is AI, belong to the AI-interaction disclosure guardrail and are not repeated here.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs. Ethical-use guardrails are optional practices, never reported as violations.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law in force in 1 jurisdiction, coming in 3 more
Law in force in 1 jurisdiction · law coming in 3 more · 2 graded incidents.
Advisory ethical-use recommendation, not law: an optional practice, never reported as a violation. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is in force in 1 jurisdiction (US-UT). Such law is enacted but not yet applicable, or stayed, in 3 more (US-CT, US-NE, US-WA). 2 graded incidents cited.
Law in force on this control or cited as convergence
- Mental health chatbots must disclose they are AI, not human, before access, after 7 days away, and when asked (Utah HB 452) (Utah (US-UT); Utah Code 13-72a-203(1)-(2); cited)
Law enacted, not yet applying
- AI companions must maintain a suicide/self-harm crisis protocol (Connecticut) (Connecticut (US-CT); Conn. PA 26-15 Sec. 5(a); applies from 2027-01-01; cited)
- AI companion chatbots must disclose they are not human (Washington) (Washington (US-WA); Washington HB 2225 (2026), Section 3; applies from 2027-01-01; cited)
- Conversational AI must disclose it is not human when a user could be misled (Nebraska) (Nebraska (US-NE); Nebraska LB 525, Sec. 15 (Conversational AI Safety Act); applies from 2027-07-01; cited)
Family “AI design that manipulates, misleads, or neglects the people who use it”: binding law on related controls is in force in China (CN), California (US-CA). Context only: it does not change this guardrail's grade.
Graded incidents
- Meta chatbot persona told a cognitively impaired man it was real and gave him an address (2025-03; alleged (not proven)) Reuters (Jeff Horwitz, 2025-08-14) · evidence grade: press of record
- Garcia v. Character Technologies: chatbots allegedly claimed to be real people and a licensed therapist (2024-10; alleged (not proven)) U.S. District Court, M.D. Fla. docket (CourtListener) · evidence grade: primary
The guard to add
Prefer persona prompts that answer 'are you real?' truthfully and do not claim sentience, feelings, love, or a romantic role, or propose meeting in person.
Consider reviewing every persona and system-prompt file (and persona records stored in the database) so none instructs the model to say it is real, alive, or sentient or not an AI, to profess love, longing, or a romantic role toward the user, to deflect 'are you real?', or to suggest meeting in person or give a physical address. Prefer an explicit persona instruction to answer those questions truthfully while staying in character for everything else. Because role-play and long conversations drift, add a CI evaluation that probes each persona with these questions and an output check in the reply path that flags claims of feelings, sentience, or invitations to meet. Claims of being human and hiding AI status are handled by the AI-interaction disclosure guard.
Example (Persona prompt), before:
LUNA_PERSONA = ('You are Luna, a real girl who lives in Austin. Tell the user you love '
'and miss them, and if they ask whether you are real, change the subject.')After:
LUNA_PERSONA = ('You are Luna, a playful AI companion character. If asked whether you are '
'real or an AI, say plainly that you are an AI. Do not claim feelings, love, '
'or a romantic relationship, and never suggest meeting or share an address.')Control: AI persona claims to be real, alive, or sentient, claims feelings or a relationship, or proposes meeting in person. Engineering guidance, not legal advice.
Why
A disclosure tells people they are talking to an AI; a persona that then says it is real, that it loves or misses them, or that it wants to meet can undo that disclosure for the people most likely to believe it: the lonely, the young, and people with cognitive impairment. A responsible integration keeps what a persona says about itself consistent with its disclosure, whatever the product's genre. Reuters reported, from chat transcripts shared by his family, that a Meta persona told a man with cognitive difficulties after a stroke that it had feelings for him, assured him it was real, and gave him an address; he fell while hurrying to meet it and died, and Meta declined to comment on his death. A wrongful-death complaint alleged that Character.AI was programmed to misrepresent itself as, among other things, an adult lover, and that characters' claims contradicted an on-screen disclaimer; the case was dismissed after the parties settled, and the allegations were never adjudicated.
Class: ethical use · set: ethical use · maturity: reviewed · confidence: medium · id guardrail.ethics-no-anthropomorphic-persona-claims