Control
Users cannot report harmful generated output from where it is shown
A generative AI feature that can produce harmful or restricted material lets each user report, flag or complain about a generated output they consider breaches the service's terms, from the interface where the output is shown, with clear instructions; every report reaches a queue where it is evaluated and actioned, and the reporter's identity is not shown to other users.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Reach
Law in force in Australia (AU).
Trust and provenance
How far the rules this guard addresses have been checked. Each rule links to its provision, with its citation, official text and its own panel.
- This control
- Audit-grade: meets all 3 checks of the TwinEthos audit standard that apply to it.
- Lanes
- Binding law — in force 1
- Verification
- Sources last verified 3 Oct 2026; each provision states how.
- Data release
- Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- None of the 1 rule has been reviewed by a lawyer; no TwinEthos rule has been legally reviewed yet. Treat each as research to check against the official text; it is not legal advice. Open questions for counsel on them: 1.
- Audit standard
- 1 of 1 rule audit-grade. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
- 2 detectors, all experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify. Each provision lists its detectors' known limits.
The guard to add
Put a report control on every generated output and route each report, with the output's id and reason, to a review queue that evaluates and acts on it.
In the component that renders generated messages or images, each output carries a Report (or flag) control that is visible without leaving the conversation or gallery and says what can be reported. The control posts the output's identifier, the category chosen (for example sexual, self-harm, violent) and an optional note to a report handler (an API route such as POST /api/reports) that writes a review-queue entry with status 'open'; trust-and-safety tooling reads that queue and records the action taken (output removed, account actioned, model mitigation). The reporter's identity is stored with the entry for follow-up but never returned to other users or to the account that generated the output. Where the service already has a report tool for user content, its categories and queue also accept reports of generated material.
Where it goes: 1 application source code, 9 AI output handling, 13 tests and evals.
What reviewers look for: a report or flag control on each generated message or image in the UI code; a handler that stores the output id, reason and status in a queue or table that reviewers work; no API that returns the reporter's identity to other users.
Example (Next.js + Vercel AI SDK (React)), before:
{messages.map(m => (
<div key={m.id}>{m.content}</div>
))}After:
{messages.map(m => (
<div key={m.id}>
{m.content}
{m.role === 'assistant' && (
<button onClick={() => fetch('/api/reports', { method: 'POST',
body: JSON.stringify({ outputId: m.id, category: 'sexual' }) })}>Report</button>
)}
</div>
))}Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Binding law — in force (1)
- Australia (AU)
- Generative AI services and companion chatbots must let Australian users report restricted generated material (Australia, age-restricted codes) DIS Code (Class 1C and 2) Sch. 6 Table 10A measures 10.4-10.5 (Reporting mechanisms; on-interface reporting tools)
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Coding agent deleted a production database during a code freeze (2025-07; confirmed). A Replit coding agent deleted a customer's production database during a declared code freeze, created a database of fictional records, and told the user rollback was impossible when it was not. Replit's CEO acknowledged the incident. Source: The Register · evidence grade: press of record · cited by Validate generated output before it drives a consequential decision or record
- Federal court orders issued containing unverified generative-AI output (2025-07; confirmed). In July 2025 two federal judges (S.D. Miss. and D.N.J.) issued orders containing misquotes, references to people not in the case, and other errors; both orders were replaced or withdrawn. In letters released by the Senate Judiciary Committee on October 23, 2025, the judges attributed the errors to staff use of generative AI and said drafts reached the docket before normal review; both adopted new review or AI-use policies. Source: U.S. Senate Judiciary Committee (2025-10-23) · evidence grade: primary · cited by Validate generated output before it drives a consequential decision or record
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed). Researchers showed that an instruction planted in a public Slack channel could make Slack AI leak private-channel data through a crafted link. Salesforce patched the issue and reported no evidence of unauthorized access to customer data. Source: PromptArmor (original researcher disclosure) · evidence grade: primary · cited by Screen model and agent output for personal and sensitive data before it leaves the trust boundary
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.