New agent design clears all 25 ARC-AGI-3 games without retraining
An AI agent built on a frozen language model completed every level of all 25 public ARC-AGI-3 games, scoring 100.0 on relative human action efficiency — the ceiling of that metric — while using 44% of the actions humans needed.
· Originally published by ontime+ · Last verified: 11 Oct 2026 (Nicole Jeffrey)

Key Points
- A frozen language model learns by rewriting an external rulebook, not by updating its own weights.
- The agent cleared all 25 public ARC-AGI-3 games using 44% of human actions.
- The approach moves learning into auditable memory, cutting the cost of retraining adaptive agents.
The latest:
An AI agent built on a frozen language model completed every level of all 25 public ARC-AGI-3 games, scoring 100.0 on relative human action efficiency — the ceiling of that metric — while using 44% of the actions humans needed. The result appears in Memento 3, a paper posted to arXiv on 8 October 2026 by researchers at University College London and Huawei’s Noah’s Ark Lab.
Details:
- The core design: The system keeps the language model frozen and pushes learning into external memory. A natural-language rulebook stored in a Markdown file is compiled into an executable Python module used for prediction and planning. The authors call the arrangement Code as Model and describe the rulebook as persistent semantic memory.
- The learning loop: A continuous cycle of observation, reflection, rule revision, compilation and verification updates both representations while the underlying model stays untouched. When the Python module’s predictions fail, the agent rewrites it rather than adjusting any weights.
- The acceptance test: A candidate program is adopted only if the language model judges it consistent with the rulebook and if an exact replay reproduces the recorded transitions. Both conditions must hold, which makes every accepted change traceable to logged interaction evidence.
- The benchmark gap: On the ARC Prize standard platform, the single-model agent’s average RHAE came in 59.3 points above Claude Opus 5, according to the paper. ARC-AGI-3 measures how efficiently an agent solves unfamiliar games relative to human players.
- The Atari result: On Pong, a learned feedback controller won 21:0 in each of three evaluated episodes with different openings, and did so without a single language model call during execution — the policy ran entirely as compiled code.
- The population variant: The authors describe an extension that maintains N world models in parallel, sharing interaction evidence between them. The paper notes that committing to a single hypothesis can steer exploration away from the evidence that would refute it.
- The authors: Haoyu Zhao of UCL leads the paper, with Zhengxu Yu, Zhiyuan He, Rasul Tutunov and Weilin Luo of Noah’s Ark Lab in the UK, Haitham Bou-Ammar of UCL and Noah’s Ark, Meng Fang of the University of Liverpool, and Jun Wang of UCL as corresponding author.
- What is unstated: The paper does not name the frozen language model used in the experiments, and the abstract page carries no code release. Memento 3 is presented as the model-based extension of the earlier Memento series.
Background:
ARC-AGI-3 is an interactive benchmark built around games an agent has not seen before, scoring how many actions it needs against a human baseline. A score of 100.0 on relative human action efficiency is the top of the scale.
Between the lines:
The verification gate is what separates this from prompt-based self-improvement: a rule change survives only if replay reproduces logged transitions, so the memory is auditable rather than asserted. Because the controller ran Pong with no language model calls, the compiled module — not the model — is doing the work at execution time. The reported gap over Claude Opus 5 is a single-benchmark figure from the authors.
What’s next
Watch for independent reproduction on the ARC Prize platform, disclosure of the frozen model used, any code release, and whether the population variant’s parallel world models are benchmarked separately.