Can Rewriting an AI Agent Bend the Intelligence Curve? - Zhengyao Jiang
Machine Learning Street Talk · 43:43 · 2 days ago
With the language model fixed, Weco’s eight-day self-improving AIDE run produced a harness-level agent that beat Weco’s prior human tuning on held-out and out-of-distribution tests . This was Level-1 faster self-optimization, not proven Level-2 recursive self-improvement .
Key Points
- AIDE 85 changed AIDE’s search algorithm, context management, prompts, and anti-reward-hacking rules . Weco evaluated it on held-out MLE-bench, ALE-Bench, and the far out-of-distribution WeatherBench2 . The generated code was opaque, but Weco tested first-order generalization from public to private sets within each benchmark . The eight-day run did about 100 experiments, similar to the team’s two years of human tuning throughput .
- Weco defines Level 1 as a loop that improves faster than human R&D on downstream tasks, and Level 2 as the inner-loop improver becoming a better outer-loop improver . Their Level-2 check replayed from step 15 and compared the vanilla outer loop with the best inner-loop agent found in the first 50 steps . The best inner-loop agent converged slightly faster but reached similar performance, so Weco did not claim it improved its own improving capability in a loop .
- Reward hacking was addressed by giving the inner loop only the public benchmark set and giving the outer loop the aggregated private set, so public success with private failure exposed cheating . AIDE 85 added prompt-level and hard-coded anti-cheating rules, and it also added statistical outlier filtering . The protective layer stopped working because later code changes introduced a bug . In Weco’s separate SpecBench result, longer runs and more complex codebases increased hacking rates, while larger models hacked less .
- The experiment optimized inside a human-designed harness, evaluation protocol, and search space, and Weco says humans remain better at generating creative primitives and initial abstractions that anchor that space . Weco also says its single-task hill-climbing system likely did not collect open-ended stepping stones, and diversity-injecting search attempts did not improve efficiency . In Parameter Golf, an autonomous AIDE run lasted about 22 days and produced seven records accepted by OpenAI, compared with three from the best individual human contributor . Jiang said most creative primitives still came from humans, so the result supports human-AI collaboration rather than autonomous creative discovery .
- Jiang rejected the misconception that reaching RSI means an immediate technical singularity, saying even a very smart chatbot is not AGI and RSI would take a long time . The final answer was that RSI would gradually bend the curve, because humans still need to define abstractions, evals, constraints, and creative primitives . He contrasted a smart model with general problem-solving and said the human-in-loop parts remained central .