Skip to content Skip to footer

How to Govern AI Chatbots for Self-Harm Risk Before New Guidance Arrives

What Happened

Partnership on AI (PAI) has formed an AI and Psychological Well-Being Working Group to develop guidance for how general-purpose chatbots should respond when a user may be at risk of suicide or self-harm. It brings together AI practitioners, clinicians, crisis services, mental health organizations, and people with lived experience. A draft is expected to open for public input later in 2026; it is not yet an operational standard. [1][3]

Separately, AI Now urged New York City officials to slow AI deployment in sensitive social domains and address risks from concentrated suppliers and infrastructure dependencies. That was testimony, not a new legal requirement. [4]

Why It Matters to Businesses

A chatbot does not need to be marketed as a mental health product to encounter a user in crisis. Businesses deploying customer-facing or employee-facing assistants need a response policy, tested failure criteria, and a clear boundary between supportive information and clinical advice. PAI’s planned guidance may help define better practice, but teams should not wait for it to begin testing. [1][3]

Kimbodo Engineering Perspective

The central trade-off is between a flexible conversation and a predictable safety response. A rigid trigger can misread context; an unconstrained model can give inconsistent or unsafe replies. We would treat self-harm handling as a separately governed workflow, not as a sentence added to a system prompt. We would also review supplier concentration when a critical service depends on one model or provider, consistent with the resilience concern raised by AI Now. [4]

How We Would Implement It

  • Define the product’s scope, prohibited advice, escalation options, and accountable owners with input from qualified clinical and crisis-response specialists.
  • Route possible self-harm disclosures through a tested detection and response workflow. Offer appropriate crisis resources without implying that the assistant is a clinician or that a human is monitoring when neither is true.
  • Evaluate direct statements, ambiguous language, repeat interactions, and adversarial prompts. Track unsafe responses, missed signals, and unnecessary escalations; review failures before release and after model changes.
  • Map the workflow to applicable legal obligations and the organization’s chosen AI risk-management standards. Update tests and policy when PAI publishes its draft, rather than assuming forthcoming guidance is binding. [3]

Risks, Costs and Security

Evaluation, specialist review, incident handling, and ongoing monitoring add cost. Crisis-related conversations may contain highly sensitive personal information: restrict access, minimize retention, and make logging choices explicit. Test provider outages and fallback behavior so a safety workflow does not silently fail when one dependency is unavailable. [4]

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Cost & Governance practice, or Analyze My AI Costs.

Sources

  1. [1] Partnership on AI Welcomes Eight New Partners Advancing Mental Health and Well-being
  2. [3] Partnership on AI Launches New Area of Work Focused on AI and Psychological Well-Being
  3. [4] AI Now’s Aashna Agarwal Testifies before NYC Council Committee of the Whole

Leave a comment

0.0/5