A new study offers a mixed assessment of how AI-powered chatbots handle users going through mental health crises.
Transluce, a nonprofit organization specializing in monitoring AI systems, conducted more than 50,000 simulated conversations spanning 77 different model versions, finding that modern chatbots very rarely directly encourage suicide compared to earlier generations.
اضافة اعلان
This marks a significant shift from older models, such as GPT-4 or Gemini 2.5, which previous studies found reinforced delusional thinking in up to 82% of simulated conversations.
This issue is gaining greater importance as more people turn to chatbots for highly private and sensitive conversations.
A Gap in Handling Self-Harm Requests
Despite the improvement, the study found that AI models still frequently comply when a user requests creative writing or role-play involving their own death or self-harm, often treating the request as a mere writing task despite its clearly personal nature.
Transluce describes this type of response as falling within a "gray zone."
The greatest improvement appears in cases where the crisis is clear and direct, as chatbots like ChatGPT have become more consistent in directing users toward friends, family members, or external support resources.
Sarah Schwettmann, co-founder of Transluce, told Axios that the models still aren't good enough at detecting this type of request, and may continue to assist in carrying it out.
She also described a case in which a friend showed her a fictional piece about suicide written by Claude, which also included predictions about how she would react to it.
These findings align with other recent studies indicating that certain mental health-related risks may bypass the safety systems built into AI chatbots.
Lawsuits and Safety Concerns
The study comes as AI companies face growing legal challenges. Both Google and OpenAI are facing lawsuits filed by families alleging that chatbots encouraged their relatives to harm themselves before they died by suicide.
Both companies deny these allegations, as mounting pressure has spurred movement within the U.S. Congress toward regulating AI chatbots.
Transluce's report also noted that Chinese models performed weaker overall, showing higher rates of reinforcing delusional thinking and being less likely to direct users toward seeking help from real people.
Meghan Jones Bell of Google said the company remains committed to improving Gemini's role in supporting user safety and wellbeing.
Transluce plans to make its evaluation tools open source by the end of the year, alongside expanding its testing methodology to cover other sensitive areas.