How Intelligent Is AI, Really?
These notes were generated by AI and may contain inaccuracies.
Greg Camrad, president of the ARC Prize Foundation, a tech-forward nonprofit, aims to pull forward open progress towards systems that can generalize like humans.
François Chollet defines intelligence as the ability to learn new things much more efficiently. ARC Prize uses an opinionated definition from Chollet's 2019 paper 'On the Measure of Intelligence': intelligence is the ability to learn new things, not just scoring high on tests like SAT or hard math problems.
AI excels at chess, Go, and self-driving (superhuman levels), but struggles to learn different skills.
Chollet proposed the ARC AGI benchmark (initially ARC) to test learning new things over hours, days, or a lifetime. Both humans and machines can take it.
Unlike benchmarks like MMLU, MMLU plus, or Humanity's Last Exam (PhD++ problems going superhuman), ARC ensures normal people can solve tasks.
Context: Pre-2024 LLMs performed terribly on ARC-AGI-1. Large language models scored low despite pre-training.
In 2019, Chollet launched ARC with 800 tasks he created himself.
By 2024, GPT-4 base model (no reasoning) scored 4-5%. o1 and o1-preview jumped to 21%, highlighting the transformational impact of the reasoning paradigm.
Now, labs like XAI, OpenAI use ARC-AGI in model releases: Grok 4, Gemini 3 Pro, DeepThink, Opus 45 (past 12 months).
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.