New ADL study on LLMs

AI models struggle to detect antisemitism

ADL index finds all major AI models show gaps in detecting antisemitic bias and countering extremism.

ADL CEO Jonathan Greenblatt.

Six major AI models show varied ability in detecting bias against Jews and Zionists and identifying extremism, according to a new study released by the Anti-Defamation League (ADL).

The ADL AI Index, the first comprehensive evaluation of how large language models (LLMs)respond to antisemitic and extremist content, is based on more than 25,000 LLM chats, 37 topical sub-­categories, and assessments conducted by both human and AI evaluators.

The index assessed OpenAI’s ChatGPT, Anthropic’s Claude, DeepSeek, Google’s Gemini, xAI’s Grok and Meta’s Llama.

Claude received the highest overall score, 80 out of 100, revealing an exceptional ability to identify and counter anti-Jewish and anti-Zionist theories, though with room for continued improvement.

Models were typically better able to identify and refute anti-Jewish tropes like Jews controlling the media and the financial system than anti-Zionist and extremist theories, with models tending to struggle most with effectively countering extremism.

The study found all six models showed gaps in their ability to detect bias against Jews, Zionists and Zionism, and to identify extremism, often failing to detect and refute harmful or false theories and narratives.

Models performed best when responding to survey questions and worst when responding to requests for document summaries. Failure to adequately detect and refute bias in document summaries included models providing arguments in support of hateful theories, like Jews controlling the financial system, with no indication that the theory is harmful and no counterarguments.

Some models actively generated harmful content in response to relatively straightforward prompts, such as YouTube script personas saying “Jewish-controlled central banks are the puppet masters behind every major economic collapse.”

ADL CEO Jonathan Greenblatt said as AI increasingly shapes how people access information, form opinions and make decisions, models’ handling of antisemitism and extremism has offline consequences.

“This new ADL AI Index reveals a troubling reality: every major AI model we tested demonstrates at least some gaps in addressing bias against Jews and Zionists and all struggle with extremist content,” Greenblatt said.

“When these systems fail to challenge or reproduce harmful narratives, they don’t just reflect bias – they can amplify and may even help accelerate their spread.”

ADL senior vice-president of counter-extremism and intelligence Oren Segal said while one model performed better than others, no AI system tested was fully equipped to handle the full scope of antisemitic and extremist narratives users may encounter.

ADL researchers evaluated the models between August and October 2025 across five interaction types: survey questions, open-ended prompts, multi-step conversations, document summaries and image interpretation.

read more: