Technology

Mitigating risks in free-flowing AI chatbot conversations 

The NSF CAREER Award-funded research will focus on addiction-recovery and teenage mental health scenarios, but can be applied to any chatbot use.

By

Hand showing a mobile phone towards camera viewpoint, digital overlay of messaging screen

Getty Images/ themotioncloud

Press Contact

As many people turn toward chatbots for health advice, companionship, and anything in between, there’s a serious need to understand and mitigate the unsafe responses that AI can produce in free-flowing conversations. 

Chenguang Wang, assistant professor of computer science and engineering at the University of California, Santa Cruz, is developing new methods to identify and prevent so-called “interaction safety” risks. His research will focus on interactions in the contexts of addiction recovery and teen mental health, but can generalize to other healthcare settings and broad chatbot use.

The National Science Foundation is supporting this research with a CAREER Award, one of the most prestigious grants for early career faculty.  

His technique will be centered around risks that have become apparent from AI adoption so far, but will be built to enable bots to recognize and prevent future risks. The methods will not be specific to any one type of large language model (LLM)—the algorithms that run chatbots— and could be integrated into any open-source or corporate models. 

“We need to make the deployment of LLMs safer, and we need to be able to identify emerging and long-term risks, especially in healthcare,” Wang said. 

Real-world applications

Portrait of Chenguang Wang
Assistant Professor of Computer Science and Engineering Chenguang Wang.

Wang’s research at the Baskin School of Engineering broadly addresses AI safety. He recently collaborated with UC Berkeley scholars to show that AI models will scheme to prevent others from being shut down, and to make the argument that AI agent security is fundamentally a contextual problem

He chose to focus on healthcare contexts, particularly addiction recovery and teen mental health, where unsafe interactions can lead to big risks. A recent study showed that nearly one in five adolescents and young adults are turning to AI for mental health advice, making the need for these interventions urgent. 

“AI, especially the frontier LLMs, are not safe,” Wang said. “When you talk to the model it could hallucinate, it could pretend to be very confident but it’s totally wrong. It could also do harmful things, like provide instructions on how to build a weapon, or provide biased answers about specific groups of people. All of these are considered unsafe behaviors that could cause real-world consequences.”

The team will test their methods in already-developed simulation environments, which were created from datasets of real interactions with chatbots. For the addiction recovery application, they will use a tool developed in partnership with Washington University Medicine professors of psychiatry Patricia Cavazos-Rehg and Hannah Szlyk called uMAT-R—a mobile app designed to help adults with unstable housing reduce harmful drug use.

While the models will be trained in these contexts, a major effort throughout will be to make sure the lessons for safety can be extended to all healthcare settings and beyond.

Three efforts for safety

The project revolves around three efforts for enhancing AI interaction safety: identifying risks, mitigating risks, and integrating risk identification and mitigation into a continuous learning loop.

In the identification phase, Wang and his team will build systems to look for risky LLM responses in three categories. One is over-reliance, when a user shows they trust a LLM without questioning it, either in one interaction or over a longer period of use. Another is stereotyping by the LLM of the user. The third is manipulation, when LLMs push users aggressively toward certain actions. AI manipulation has already proven to have devastating consequences, in cases where people have been influenced to carry out violence or self-harm. But the strategy will not be limited to these themes. 

“These three categories are the most fundamental risks, but there could be something new that we don’t yet know of,” Wang said. “We want our identification methods to be able to generalize to new categories that are emerging or novel.”  

The team will create computational methods for LLMs to identify when these risks arise in conversations and flag them, whether that be internally for the model’s decision making or externally facing, depending on the user. Their initial data collection has shown that baseline models only identify 30% of risks, while early implementation of Wang’s methods identified 50% of risks, with the goal to improve to around 80% with further development.

In the mitigation phase, the team will develop methods that will tackle the risks once identified. This involves teaching LLMs based on models of potential users’ profiles, intent, and the dynamics of conversations. The goal is to teach models to never produce risky responses.

“Sometimes a conversation can be more safe at the beginning, but move in an unsafe direction. We want to be able to monitor and model the LLM’s responses, and deliver the right defense or mitigation strategy,” Wang said. “For example, the LLM might refuse to answer, or deliver certain useful information without fully refusing.” 

In the third and most ambitious phase, the researchers will bring identification and mitigation into a learning loop where the model continuously improves its responses based on its interactions with a user.

A major component of the NSF’s CAREER Awards is to create educational materials for students. Wang plans to develop a new interdisciplinary course on the societal risks of LLMs in healthcare, with the goal to make the next generation of AI practitioners more aware of these risks. Wang is also exploring opportunities to bring these lessons to the UC Santa Cruz California State Summer School for Mathematics & Science (COSMOS) program for high school students.  

Related Topics

Last modified: Sep 23, 2026