AIDE2

Type: types/note.md

Evidence basis: Recursive self-improvement of AI research agents, captured 2026-09-24; doc-grounded analysis of the described two-loop improvement plane, without agent source or original run artifacts.

AIDE2 searches for better research-agent harnesses. An inner agent edits task code using public feedback. The outer agent rewrites the current research-agent code, grades each candidate on returned solutions' hidden task scores under fixed budgets, and retains the best agent. The aggregate private grade is visible to outer search, so further external benchmarks provide a separate generalization boundary.

The paper reports seven accepted rewrites in its main eight-day run and improved performance on four held-out benchmarks. The evolved AIDE85 changes both search allocation and context management: strategy-arm search replaces greedy expansion, bounded root/recent-candidate context replaces growing full history, and recurring error lines enter prompts when the bug rate reaches 15%. Accepted harness code persists across tasks; the exact storage and persistence of compact summaries are not established by the paper.

Several distinctions constrain interpretation. A robustness penalty reportedly never changed final candidate choice, so presence in the winning agent does not establish contribution. A reported patch also changes a held-out evaluator's failure handling; its independent admission and measurement-integrity boundary remain uninspected. Reduced reward hacking is reported for the evolved bundle, without identifying the responsible rewrite. The ignition experiment shows a discovered agent can drive further improvement but does not establish that it is a better self-improver than the human-engineered outer agent.

Scope

The exact retained analysis contains source anchors, route and authority records, both mandatory lenses and all fourteen memory-comparison axes. Self-improvement and reflection are claimed at the described harness-lineage boundary. Criticism-specific conjectural learning, exact model fixation, implementation enforcement and independently confirmed component effects remain unestablished. The reported outcomes support a bounded improvement finding; they do not establish accelerating self-improvement or a general guarantee against metric exploitation.