Moravec’s Paradox and AI Benchmarks

AI can excel at tasks people find difficult while struggling with tasks that seem easy. FrontierMath highlights a similar gap in mathematical reasoning and raises broader questions about how AI should be evaluated.

Moravec’s Paradox

When comparing artificial intelligence with human abilities, the differences often appear in ways that run counter to intuition. In the 1970s, the American roboticist Hans Moravec described an interesting paradox in AI research: AI can perform tasks that humans consider difficult, such as playing chess, yet finds tasks that are easy for humans, such as tying shoelaces, extremely difficult.

This is known as Moravec’s paradox. It arises from the difference between human intuitive understanding and a machine’s computational approach.

Related Lancet Correspondence: https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(23)01129-7/fulltext

Why Better Benchmarks Are Needed

A recently released benchmark from Epoch AI, FrontierMath, makes this paradox more apparent. FrontierMath is intended to assess mathematical reasoning by presenting current AI models with extremely difficult mathematics problems. Models that achieved high scores on earlier mathematics benchmarks, such as GSM8K and MATH, reportedly achieve only around 2% on FrontierMath.

Existing benchmarks are often described as saturated. In part, large language models have evolved to score well on benchmarks, and there is also the problem of contamination: answers may already be present in the training data.

Moravec’s paradox explains why tasks that appear easy to people may be difficult for AI. Highly structured games such as chess have clear rules and finite state spaces, allowing AI to solve them quickly and efficiently. Tying shoelaces, by contrast, requires physical sensation, fine manipulation, and spatial awareness, making it a much harder task for AI.

A similar pattern appears in mathematical problem-solving. AI may perform exceptionally within a particular range of problems, while its limitations become visible on problems such as those in FrontierMath. FrontierMath spans all major areas of mathematics and consists of problems not included in existing datasets, so it does not have the same data-contamination issue. Its answers are also designed to be difficult to obtain by guessing, meaning that a model must actually understand and solve the problem logically to succeed.

The Difficulty of FrontierMath

FrontierMath makes the difference between human and AI thinking more visible. When models that exceed 90% success on existing mathematics benchmarks score below 2% on FrontierMath, this suggests that AI can solve certain types of problems well but has difficulty with tasks requiring long-context maintenance and creative problem-solving.

For this reason, FrontierMath is attracting attention as a tool for evaluating AI’s real capabilities. Human mathematical thinking requires more than calculations performed through prescribed steps; it also requires creative ideas and refined reasoning. FrontierMath demands precisely these abilities from AI. Because it calls for a new mode of thinking rather than standardized problem-solving, its low success rate should not be seen merely as failure, but as an important compass for AI research.

The Difficulty of What Seems “Easy”

As Moravec’s paradox suggests, the fact that AI can solve complex problems does not by itself mean that it has reached human-level thinking. High-difficulty benchmarks such as FrontierMath are important, but it is equally important to evaluate how AI performs everyday, seemingly easy tasks.

New benchmarks are needed to assess abilities that are easy for humans, such as maintaining long-term context, solving problems autonomously, and developing a consistent line of reasoning. If such evaluation measures are established, AI may develop beyond a tool that solves particular problems and become a genuine assistant that can collaborate with people.

This may be easiest to understand by considering cases in which people working in AI are astonished by an achievement that the general public regards with indifference, or the reverse.

Terence Tao’s assessment is also striking: he suggests that AI will struggle for the next few years. Perhaps humans can hold it back for only a few years, after which we may enter an era in which humans can no longer evaluate AI at all.

Source

Moravec's paradox in LLM evals

“I was reacting to this new benchmark of frontier math where LLMs only solve 2%. It was introduced because LLMs are increasingly crushing existing math benchmarks. The interesting issue is that even though by many accounts (/evals), LLMs are inching…” — Andrej Karpathy, November 10, 2024: https://t.co/3Ebm7MWX1G