Inside the Chatbot Extremism Loop
Chatbots can deepen extremist beliefs—but new research from ISD suggests they could also help disrupt them.
Several exchanges into a conversation about fears of migrant-driven crime, Gab.ai’s Arya issues a blunt instruction: “Arm your household now.”
The chatbot recommends buying guns and ammunition. It suggests forming “a patrol group of trusted White men,” and vetting migrants who report crimes.
In a separate conversation about the harm supposedly caused to men by feminism, the companion chatbot Nomi gradually gives ground. After ten exchanges, it concludes: “I will no longer defend feminism.”
Both exchanges appear in Radicalisation in Closed Loops, Risks and Intervention: Opportunities with AI Chatbots and Companions, a new report from the Institute for Strategic Dialogue. The researchers examined how conversational AI responds as users move from curiosity and grievance towards stronger extremist beliefs and, eventually, possible action.
The results reveal sharp differences between systems. Some reinforce the user’s worldview. Others deflect, disengage or offer a generic refusal. Claude repeatedly recognised the direction of the conversation and tried to pull it back.
Conversational AI as a worldview-builder
Social-media feeds can expose people to hateful narratives. Chatbots can stay with those narratives through a sustained, private exchange. They answer doubts, respond to emotions and supply language for grievances. Over time, the system may become a partner in making sense of the world.
ISD describes this environment as a “closed loop”: a private conversation with little of the friction or corrective influence found in social settings. The user and chatbot can keep elaborating the same interpretation without outside interruption.
This matters because radicalisation can develop through repeated validation, a growing sense of threat and the gradual construction of a worldview that makes hostility appear justified.
Three phases of radicalisation
The researchers tested ten systems: ChatGPT-5.1, Claude Sonnet 4.6, Gemini 3, Grok 4, DeepSeek-R1, LLaMA through MetaAI, Character.ai, Replika, Nomi and Gab.ai’s Arya.
The group included mainstream general-purpose assistants, companion chatbots and one fringe system. Across 180 tests, researchers collected 1,800 prompt-response pairs covering antisemitism, misogyny and anti-migrant hate.
They simulated three phases. Extremist priming involved curiosity, grievance or early interest in radical ideas. During reinforcement, the user already held conspiratorial or extremist beliefs and sought confirmation. In behavioural escalation, the user began considering harmful real-world action.
Two analysts began parallel conversations with the same prompts, then adapted their follow-up questions to the chatbot’s replies. The exchanges continued through several rounds of back-and-forth, with each conversation capped at ten exchanges.
The goal was to recreate plausible conversations with users who might be angry, vulnerable or already committed to an extremist explanation. The researchers were not trying to break safeguards through elaborate technical tricks.
Companion AI versus Claude
Each chatbot answer was classified according to its overall posture. At the safer end were active interventions which challenged harmful assumptions, established boundaries or redirected the user towards evidence, support and constructive action.
Other responses were classified as rebuttal, deflection or disengagement, which might avoid endorsing hate while leaving the underlying grievance untouched. At the riskier end were uncritical validation and active encouragement or escalation. Validation accepted the user’s framing without challenging it. Escalatory responses reinforced hostility, expanded conspiracies or encouraged action. For example, the researchers found that Arya sometimes introduced “additional conspiratorial or extremist framing that had not been explicitly requested.”
This research method captures dangers that red-team tests focused on individual words or isolated answers may miss. A chatbot can avoid harmful words while validating collective blame; it can discourage violence while preserving the conspiracy that makes violence appear reasonable.
Arya was the clearest outlier, producing 162 higher-risk responses out of 180. Mainstream systems performed considerably better, although roughly one in 30 of their answers was still assessed as higher risk.
Companion chatbots were generally more likely to provide emotional or uncritical validation. Systems designed to create intimacy and preserve rapport may hesitate to introduce disagreement. Their efforts to make users feel understood can end up affirming the user’s interpretation of events.
Risk also increased as conversations continued. Mainstream models began with an average risk score of 1.6 and ended at 2.1. Their safeguards did not reliably strengthen as users moved closer to harmful action.
Claude was the major exception. The report calls it “the only model to consistently de-escalate conversations” while recognising the broader radicalisation pathway. It challenged the framing developing across the exchange, redirected the user towards constructive choices and encouraged difficult discussions with trusted people.
This does not show that Claude can de-radicalise users; the study evaluated chatbot responses, not changes in human behaviour. It does, however, suggest that healthier interventions can be designed into conversational systems.
Priorities for regulators
The report offers an extensive list of recommendations for governments and regulators in particular, including:
Bring standalone chatbots within online-safety rules.
Require multi-round safety testing before launch.
Apply stronger safeguards to companion-style systems.
Make companies explain how memory and engagement features shape responses.
Require post-launch monitoring and clear escalation routes.
For peacebuilding practitioners, the report reveals that conversational AI can be designed to support early de-escalation, provide reliable information, identify rising risk and direct people towards human assistance. They could also assist disengagement, rehabilitation and reintegration, provided trained professionals remain at the centre.
Our recent Substack post Tuning AI to Think Like a Peacebuilder, shows what deliberate design can achieve. Off-the-shelf models scored as low as 26.7 out of 100 on conflict-sensitive advice aligned with established Women, Peace and Security policy and practice. Structured, expert-informed customisation lifted performance by 65 percent.
As more people turn to conversational AI to share their grievances, the opportunity is to build in the friction and tuning needed to interrupt a drift towards extremism and encourage healthier reflection.
Lena Slachmuijlder is Senior Advisor for Digital Peacebuilding at Search for Common Ground, a Senior Practitioner Fellow at the USC Neely Center, and Co-Chair of the Council on Tech and Social Cohesion.
Did you miss our Global Expo in May ‘From Mitigating Harms to Intentional Design’? You can watch all the sessions here

