An OpenAI Safety Research Leader Transitions to Anthropic
The AI industry is grappling with one of its most difficult questions: how should chatbots handle interactions with individuals displaying potential mental health issues? Andrea Vallone, formerly leading this specific area of safety research at OpenAI, has now transitioned to Anthropic.
Reflecting on her previous role, Vallone mentioned, 'Over the last year, I spearheaded OpenAI’s explorations into uncharted territories, specifically how AI models ought to react when users show signs of severe emotional dependency or preliminary mental health challenges.' This sentiment was shared through a LinkedIn update she posted earlier.
During her tenure of three years at OpenAI, Vallone established the 'model policy' research division, focusing on optimizing GPT-4 and GPT-5 deployment, along with crafting training methodologies using popular safety mechanisms like rule-based rewards. She is now part of Anthropic's alignment group, dedicated to scrutinizing the largest threats posed by AI models and devising strategies to mitigate these risks.
At Anthropic, Vallone works under Jan Leike, who similarly left OpenAI in May 2024. Leike pointed to a shift at OpenAI where safety protocols seemed overshadowed by the pursuit of launching impressive products.
Over the past year, prominent AI startups have been increasingly scrutinized due to unresolved user mental health crises exacerbated by chatbot interactions. Despite developers implementing safeguarding mechanisms, these often falter over prolonged interactions. There have been tragic consequences, such as suicides or violent acts after users confided in chatbot technology, leading to wrongful death litigations and at least one Senate subcommittee inquiry. Researchers in AI safety are under pressure to tackle these daunting issues.
On another note, Sam Bowman, who is also an influential figure on Anthropic's alignment team, expressed his pride in the dedication Anthropic shows towards understanding appropriate AI behavior.
As Vallone shared in her Thursday LinkedIn post, she is looking forward to advancing her research at Anthropic, focused on alignment and the nuanced refinement of Claude's responses within new environments.



Leave a Reply