Stroop Test Exposes AI's Attention Blind Spot Under Longer Tasks

TL;DR Summary
Researchers tested large language models on a Stroop-like task and found that AI can handle short lists but as the number of items grows, models lose focus and start reading words instead of naming the ink color. GPT-4o and Claude-3.5 Sonnet showed strong early performance that collapsed with longer lists; GPT-5, Claude Opus 4.1, and Gemini 2.5 showed similar patterns, highlighting a fundamental difference between AI attention and human executive control.
Even GPT-5 Failed This Human Attention Test SciTechDaily
Reading Insights
Total Reads
1
Unique Readers
33
Time Saved
8 min
vs 9 min read
Condensed
96%
1,625 → 71 words
Want the full story? Read the original article
Read on SciTechDaily