Advanced AI models suffer a near-total collapse on classic psychology test as cognitive demands increase
Leading AI models experience a catastrophic performance collapse on the Stroop task. Accuracy falls as tasks grow longer and more complex. This suggests these systems lack human-like executive control.
What changed
No new corroborating evidence was provided regarding the specific PNAS Nexus study results.
Live updates
-
AI Models Fail Stroop Task as Complexity Increases
confidence 100%Leading AI models experience a catastrophic performance collapse on the Stroop task. Accuracy falls as tasks grow longer and more complex. This suggests these systems lack human-like executive control.
What's confirmed:
- AI models show a catastrophic performance collapse on the Stroop task.
- Accuracy in these models drops as tasks become longer and more complex.
- The performance drop indicates a lack of human-like executive control.
-
Advanced AI Models Fail Classic Psychology Attention Test
confidence 90%A study published in PNAS Nexus reveals that leading AI models suffer a catastrophic performance collapse on the Stroop task. While systems can identify colors in short lists, accuracy drops sharply as tasks become longer and more complex. Researchers suggest this indicates a lack of human-like executive control.
What's confirmed:
- Leading AI models dropped from over 90% accuracy to near-zero or complete failure as the complexity of the Stroop task increased.
- The study was published in PNAS Nexus.
- The Stroop test requires participants to identify the color of a word's text while ignoring the word itself.
- Research indicates AI systems struggle with sustained focus and conflict resolution compared to human attention.
Still unconfirmed:
- GPT-5 failed the human attention test.
- The findings raise concerns for autonomous drone navigation and real-time obstacle avoidance.
- The lack of executive control could stall the development of human-level AI.
-
Advanced AI Models Fail Classic Psychology Attention Test
confidence 90%A PNAS Nexus study using the Stroop task revealed a catastrophic collapse in AI executive control. While models handle short lists, accuracy drops sharply as tasks grow longer and more complex. These results suggest AI lacks the human-like focus needed for artificial general intelligence.
What's confirmed:
- A study published in PNAS Nexus used the Stroop task to examine AI attention and executive control.
- AI models can correctly name colors in short lists but experience a performance collapse as task complexity increases.
- Some leading AI systems saw accuracy drop from over 90% to near-zero or near-complete failure.
- The Stroop test requires participants to identify the color of a word's text while ignoring the word itself.
- GPT-5 failed the human attention test.
Still unconfirmed:
- The failure in sustained attention may pose risks for autonomous drone navigation and real-time obstacle avoidance.
- Suketu Patel and colleagues led the research into transformer-based machine attention.