https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-...
There you will find the extremely important qualifier it's the public set, not the private set (with the risk of overfitting, ie the results not repeating when submitted to be run in competition), and the detail that this is essentially a harness added to Opus 5, not Nvidia's own models.
Obviously still impressive, you would think.
True as that may be, it may be better to optimize models for some amount of memory versus forcing some token count based on a reasoning level, right?
Most of this doesn't discredit your overall point, though.
Community maintained spreadsheet of the runs: https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-P...
What NVIDIA has here is a generic "evolution" harness, which can be used for any problem.
I think it would be fair game to allow OpenClaw, Hermes, Codex, Grok Bot, this NVIDIA thing, to compete, as long as they don't have ARC-AGI specific skills, toolset.
GPT-4 was decidedly not capable of beating Pokemon 18 months ago. I doubt it would be able to complete a single level. I don't think people realize how large the advances in model capabilities have been. GPT-4 in a modern harness is absolutely horrendous compared to modern models.
Have you ever played pokemon?
Using Claude Opus 5, but it can use others:
AVO is also designed to operate across frontier models. While our full public-set result used Claude Opus 5, we additionally paired AVO with GPT-5.6 Sol on a challenging subset of games. In these limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions in matched-level comparisons. These preliminary results suggest complementary operating profiles across models, and we leave a broader systematic comparison to future work
Curious to read more about it though, seems the paper for it is here: https://arxiv.org/pdf/2603.24517, I'm not sure I understand if it's better than just Codex with a /goal, as they talk about "can discover performance-critical micro-architectural optimizations" but leave Codex alone for a day or two and you'll get the same results without doing "additional autonomous adaptation" at all.
If not, is the benchmark just incorrectly named? (I personally think so.)
PS: I follow Wikipedia's definition for AGI (https://en.wikipedia.org/wiki/Artificial_general_intelligenc...), which also talks about some tests. However, I distinguish it from Strong AI.
no? once 3 is solved, we would come up with 4. then 5, 6...
it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.
The approach feels asymmetric though now, since by the time AI reaches a point where humans cannot come up with any task that's easy for humans but hard for machines, it (AI) may be able to do some tasks that an average human can't, or do some much better than an average human.
No.
AGI is impossible without a biological pineal gland. The pineal gland is the seat of consciousness and without one, any AI is merely a pattern matcher, not intelligent.
(1) Consciousness is mandatory for intelligence.
(2) 'AGI' is no different than just 'intelligence'.
(3) Consciousness resides in pineal gland.
(4) Biology is mandatory for consciousness/AGI.
I cannot claim these to be wrong, however, have no reason to believe in any.
We do seem to agree though that the benchmark is incorrectly named.
Anyway yes I think we've had AGI for a while now, even if the GI doesn't quite match up with what we expect from a human.
Yey, AGI is finally solved.
100% is some "RHAE" metric: its performance of median human first time seeing those problem.
The next year is going to be wild folks