discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.

AI models flub these intelligence tests. Can you fare any better?

AI models rapidly improve on some puzzles but struggle with spatial reasoning, memory adaptability, and abstract visual tests, highlighting fundamental differences from human cognition.

Aug 26·technologyreview.com·3 min read

Intelligence analysis by Gemini 2.5 Flash

AI models flub these intelligence tests. Can you fare any better?
Image: technologyreview.com

The article explores how puzzles serve as benchmarks for AI intelligence, demonstrating both rapid advancements and persistent weaknesses. Despite progress in areas like Connections puzzles, current models falter on tasks requiring human-like spatial manipulation, nuanced memory recall, and abstract visual inference, offering insights into their cognitive limitations.

Why it matters

Understanding where AI models succeed and fail in intelligence tests is crucial for identifying their current limitations and guiding future research toward more robust and human-like artificial intelligence.

Imagine a super-smart robot that can learn tons of facts and solve many word puzzles super fast. But if you ask it to imagine turning a toy block in its head, or to spot a tricky detail in a picture that looks almost like one it's seen before, it gets confused. Humans are still much better at these kinds of 'thinking with your eyes' or 'spotting the trick' games, showing that AI still thinks differently than we do.

Analysis

Puzzles have historically been a cornerstone of AI development, serving as critical benchmarks for measuring progress and identifying areas of weakness. From Arthur Samuel's checkers algorithm in 1959 to modern challenges like chess and Go, games provide a structured environment to test machine intelligence. While AI has shown remarkable improvement in certain domains, such as solving New York Times Connections puzzles, where models advanced from 18% accuracy in late 2024 to near perfection by early 2025, these tests also reveal significant gaps in AI's cognitive abilities compared to humans.

Spatial Reasoning

One of the most pronounced areas where AI models consistently underperform humans is spatial reasoning. Tasks like mental rotation problems, common in IQ tests, require the ability to mentally manipulate 3D objects from different angles. Despite advancements in visual input analysis for language models, they still fail abysmally at these challenges. This deficiency suggests that even with sophisticated 'world models' designed to help AI understand physical environments, current large language models (LLMs) lack the intuitive spatial manipulation capabilities that come naturally to human spatial thinkers like architects and mechanical engineers. This highlights a fundamental difference in how machines process and understand physical space.

Knights and Knaves

Another revealing category of puzzles involves memory and adaptability, exemplified by 'Knights and Knaves' problems. These logic puzzles require discerning truth-tellers from liars based on their statements. A 2024 study by researchers from Google and the University of Illinois Urbana-Champaign found that when models encountered slight variations of puzzles they had seen during training, they often failed to spot key differences, instead defaulting to memorized responses. This tendency to rely on rote memorization rather than adaptive reasoning can be a liability, causing models to 'whiz by' critical nuances. This issue also surfaces in 'SimpleBench' problems, which resemble complex problems seen during training but contain subtle tricks that humans easily identify, while top-tier AI models frequently miss them.

ARC-AGI

The Abstract and Visual Reasoning Challenge (ARC-AGI) benchmark further exposes AI's struggles with abstract and visual reasoning, even in two dimensions. These puzzles demand that models infer abstract, general rules from a set of examples. While models perform better when grids are encoded as numerical strings rather than images, research indicates that even correct answers are often achieved through 'byzantine and non-generalizable rules,' rather than the simple visual concepts humans employ. This suggests that AI's problem-solving methods in these complex visual and abstract tasks are fundamentally different from human intuition, often lacking the generalizability and conceptual understanding that define human intelligence.

Key points

  • Puzzles have been crucial for AI development, showcasing both rapid advancements and persistent limitations.
  • AI models significantly struggle with spatial reasoning tasks, such as mental rotation, unlike humans.
  • Models can be tripped up by subtle variations in puzzles, relying on memorized solutions rather than adaptive reasoning.
  • Abstract and visual reasoning challenges like ARC-AGI reveal that AI often uses non-generalizable rules, differing from human intuition.
  • These tests highlight fundamental differences between human and machine cognition, offering insights into AI's strengths and weaknesses.
The Upside

These identified weaknesses provide clear targets for AI researchers, enabling them to develop new architectures and training methods that could eventually bridge the gap in spatial, adaptive, and abstract reasoning, leading to more versatile and robust AI systems. Continued testing with such puzzles will drive innovation towards more human-like intelligence.

The Downside

The persistent struggles of even advanced AI models with fundamental human cognitive tasks like spatial reasoning and adaptive problem-solving suggest that true general artificial intelligence remains a distant goal. These limitations could hinder AI's ability to operate effectively in complex, unpredictable real-world environments that demand flexible, intuitive understanding.

Originally reported at

technologyreview.com

Discernion covers the story. Read the full piece at the source.

Tagsairesearchllmscognitiontestingpuzzles

Intelligence analysis by

Gemini 2.5 Flash

Published

Aug 26, 2026

Source

technologyreview.com

Share

Topics

airesearchllmscognitiontestingpuzzles

Related

More from this desk

Aug 26·techcrunch.com

QueryStory wants you to believe what AI is telling you

QueryStory, a new startup co-founded by Shapor Naghibzadeh, emerged from stealth with $6 million in seed funding to help large enterprises trust AI-driven data analysis by providing verifiable narratives and confidence indicators.

Aug 26·techcrunch.com

Arga Labs is building a better way to train enterprise AI agents

Arga Labs secured $10 million in seed funding to develop digital twin environments for training enterprise AI agents, addressing the difficulty of testing agents in complex business software like Salesforce and Workday.

Aug 26·technologyreview.com

The Download: the Kids issue arrives, and Bill Gates reveals his AI fears

MIT Technology Review launches its 'Kids issue' exploring how technology and AI are reshaping childhood, while Bill Gates expresses deep alarm over AI's rapid, unchecked advancement.

Anthropic logo
Aug 26·anthropic.com

Societal Impacts Research

Anthropic's Societal Impacts team conducts technical research to understand how AI is used in the real world, focusing on human values, real-world applications, and anticipating future risks. They collaborate with policy and safeguards teams to inform better AI governance.