A Survey · 3DAgentWorld
When the Improver Becomes the Improvee
The survey in one paragraph pair. Read the full text for the complete version.
Learning systems increasingly help build and improve other learning systems. Searchers tune training settings, meta-learners shape learning rules, and agents edit code and workflows. Models also generate data and assess other models. We review these methods under the name AI for AI (AI4AI): the use of learned components or automated search to guide changes to AI systems and the artifacts used to build and assess them.
We trace links among reinforcement learning, automated machine learning, meta-learning, self-play, and language-model agents. A four-axis taxonomy describes lifecycle stage, improvement target, automation degree, and method family. We use it to compare what changes in each system and how feedback guides those changes. We distinguish task-level refinement, lasting self-improvement, and recursive self-improvement, which also requires a better capacity to make further improvements. We close with research priorities in independent evaluation, full cost reporting, human oversight, and controlled self-modification.
Continue reading →Four moves that turn a scattered literature into one measurable family.
Traces six research strands that overlap in time rather than forming stages: agents and RL, AutoML and meta-learning, self-play, and language-model agents. Recursive self-improvement sits at the high-automation end of the same family.
A coordinate system for comparison: lifecycle stage, target of improvement, degree of automation, and method family. One cell per axis, or a trajectory across cells over time.
Landmark methods reviewed at mechanism level: what the outer loop proposes, what the inner loop scores, which artifact changes, and whether the proposer can itself be rewritten.
Research priorities in independent evaluation, full cost reporting, human oversight, and controlled self-modification. Progress is judged by verified gains under stated conditions, not by the amount of automation.
The strands overlap in time and developed in parallel; the grouping collects ideas, not dates. See Figure 2 and Section 2 for the full account.
Nine sections, from historical grounding to an engineering agenda for self-improvement.
If this survey is useful to your work, please cite it as follows.
@article{ye2026ai4ai,
title = {A Survey on AI for AI},
author = {Ye, Deheng and Zhang, Zheng and Wang, Hao and Miao, Chunyan},
note = {Ye and Zhang contributed equally.},
year = {2026},
month = sep,
url = {https://3dagentworld.github.io/AI4AI-survey/}
}