OpenAI, DeepMind, and Anthropic Teams Race to Respond to the ARC-AGI-3 Challenge
The three major labs have announced their plans to address the ARC-AGI-3 benchmark, which measures abstract reasoning ability, amid expectations that the test will reveal real gaps between systems.

The ARC Prize organization published the third version of the ARC-AGI benchmark in August 2025, announcing that the new test fundamentally raises the bar compared to the previous two versions, as it focuses on models' ability to grasp abstract patterns not present in training data and apply them to entirely different situations.
Major labs responded with remarkable speed; the OpenAI team announced it is working on adapting the o3 model to handle the new task structure, while the DeepMind team revealed that Gemini Ultra is currently undergoing intensive internal evaluation prior to any official announcement. Anthropic researchers explained that Claude 4 is being re-tested against the updated benchmark with a focus on generalization tasks.
Challenge organizers believe ARC-AGI-3 measures what they call "first-principles reasoning"—the model's ability to build an inferential foundation from very few examples, which is considered a cornerstone in definitions of general intelligence. Benchmark developers noted that top models do not exceed 30% on its hardest tasks.
The competition over ARC-AGI-3 reveals a pivotal equation: Labs leading this benchmark will hold a powerful marketing argument in the race toward general intelligence, making every percentage point in its results strategically valuable beyond its abstract technical worth.
What do these terms mean?
ARC-AGI: Short for Abstraction and Reasoning Corpus, tests that measure AI's ability to reason abstractly and solve problems it was not trained on.
Generalization: A model's ability to apply what it has learned to new situations, like a student applying a mathematical rule to problems they haven't memorized.
First-Principles Reasoning: An approach starting from fundamental facts rather than memorized patterns, serving as a hallmark of true understanding.
Weekly Newsletter
Read between the lines before everyone else. Decode the most important economic, tech, and decision-maker movements in the region.. in 5 minutes every Saturday.
Recommended
Why Does the Academic Fail When Entering Business?
From One Person to an Institution: The Real Journey of the Family Business
أسعار الذهب ٢٠٢٦: اللي بيطبع الدولار… بيشتري الذهب
Tax Sukuk… A Brilliant Idea, But Let's Hope We Aren't Spending Tomorrow's Taxes Today
Why the Energy Grid — Not Generation — Is the Real Bottleneck
60 Heroes
.png?alt=media&token=5eaf4867-f044-4dbd-bda6-e0fef52a50df)


.png?alt=media&token=ed68f44d-e5db-4b1a-947a-42f608a0d791)
.png?alt=media&token=8c404866-0244-4cf8-9238-2e85c3b88fed)

.png?alt=media&token=539bd907-5945-4421-a543-a3fe6017c859)
.png?alt=media&token=5eaf4867-f044-4dbd-bda6-e0fef52a50df)


.png?alt=media&token=ed68f44d-e5db-4b1a-947a-42f608a0d791)
.png?alt=media&token=8c404866-0244-4cf8-9238-2e85c3b88fed)

.png?alt=media&token=539bd907-5945-4421-a543-a3fe6017c859)

