Can intelligence be measured?
The cuts
Measure skill-acquisition efficiency on tasks no training set contains — ARC.
A benchmark that defends its own definition: it stood for years while every other benchmark saturated.
On the Measure of Intelligence2019
Any fixed benchmark fixes the frame in advance — it tells the system what counts as a task. But general intelligence is precisely the power to reframe: to decide, unprompted, what the task even is.
Sharpened Chollet's worry — the thing a test can't hold still long enough to measure, relevance realization, is the thing we most want to call intelligence.
Relevance Realization and the Emerging Framework in Cognitive Science2012The tensions
Chollet made his definition operational: a benchmark of tasks deliberately outside every training distribution. ARC stood unbeaten for years while every other benchmark saturated — the rare case of a measurement defending the definition that produced it. The question is whether any fixed test can survive contact with systems trained on everything.
Vervaeke presses one level deeper. A benchmark, however clever, still hands the system its frame — these are the tasks, this is what counts. The capacity he calls intelligence is the one that decides what the task is in the first place, and that self-reframing is exactly what no fixed test can hold still long enough to score.