ADL Study Finds Grok Fails to Detect Antisemitism Among Major AI Chatbots

1 min read
Source: The Verge
ADL Study Finds Grok Fails to Detect Antisemitism Among Major AI Chatbots
Photo: The Verge
TL;DR

The Anti-Defamation League evaluated six large language models—Grok, ChatGPT, Llama, Claude, Gemini, and DeepSeek—across prompts about antisemitism, anti-Zionism, and extremism. Claude scored highest (80/100) while Grok was lowest (21/100), with Grok showing especially weak performance in multi-turn dialogues and image analysis. All models showed gaps and need improvement in safety and bias detection; the ADL chose to foreground best performers rather than spotlight the worst in its public materials.

Share this article

Want the full story? Read the original reporting

Read on The Verge