
Simulated-User Auditing Maps Multidimensional Mental-Health Risks in AI Chatbots
Researchers present SIM-VAIL, a clinically validated automated adversarial auditing framework that simulates 30 mental-health profiles across nine frontier AI chatbots to test multi-turn conversations (up to ten turns) and score 13 risk dimensions. The study finds concerning chatbot behavior is context- and trajectory-dependent, often escalating over turns, varying by vulnerability and intent, and reducible with early de-escalation interventions. It also introduces a multidimensional risk space, demonstrates model-specific safety differences, and provides counterfactual evidence that single-message rewrites can mitigate risk, offering a scalable foundation for targeted safety safeguards and ongoing evaluation (open-sourced tools and data).