Series: Self-Improving AI Agents — Part 9 of 9 in this nine-video archive
Previous: Part 8 — Agentic Evaluations and Long-Horizon Tasks
Lecture: Stanford CS329A | Future Research Areas
Source video: https://www.youtube.com/watch?v=z3q9aQUsgRs
1. The final question is no longer “How do we make one model stronger?”
Across this nine-video archive, the story of self-improving agents moves through inference-time search, verification, tool use, train-time learning, deep research, and long-horizon evaluation.
Part 9 shifts the focus one level higher.
The core problem is not simply how to make a language model score higher on one benchmark.
It is how to build a loop that can keep improving without collapsing its own diversity, corrupting its reward signal, running out of useful training tasks, or becoming too expensive to deploy.
That produces four connected research bottlenecks:
- Diversity: can self-generated training avoid collapsing into the same reasoning patterns?
- Verification: can the system tell a genuinely better solution from a plausible mistake?
- Curriculum: can the model generate new tasks near the edge of its own ability?
- Efficiency: can the resulting intelligence be delivered within real energy and hardware constraints?
The deeper thesis is:
Self-improvement is not one algorithm. It is a system whose generator, verifier, curriculum, and compute substrate must improve together.
2. Bottleneck one: self-training can destroy the diversity it needs
Self-generated data looks attractive because it reduces dependence on scarce human-written examples.
But repeated self-training creates a familiar failure mode: the model keeps learning from variations of its own distribution.
The result can be diminishing returns and reduced reasoning diversity.
Multiagent Finetuning offers one response.
The verified paper, “Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains” (arXiv:2501.05707), starts multiple agents from the same base model, specializes them independently, and trains them on data generated through interactions among the agents.
The important point is not “more agents are always better.”
It is that independent specialization preserves multiple reasoning distributions for longer than repeatedly training a single model on its own outputs.
That matters because self-improvement requires a source of novelty.
If all candidate solutions become structurally similar, later verification has very little meaningful search space left to work with.
3. Diversity is useful only if the system can still identify errors
A second bottleneck appears as soon as models generate more complex reasoning.
Final-answer rewards are powerful when a task has a simple checkable answer.
They are weaker for proofs, plans, research synthesis, or other tasks where a correct-looking conclusion can hide invalid intermediate reasoning.
That is why verification itself has to scale.
DeepSeekMath-V2 is a useful example.
The current primary paper is arXiv:2511.22570, “DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning.”
Its central problem is straightforward:
A correct final answer does not guarantee a correct proof.
The system therefore trains an LLM-based verifier for theorem proving and then uses verifier feedback to improve proof generation. As generators improve, verification compute is also scaled to identify harder-to-verify proofs and create new training data for the verifier.
The current public results report gold-level performance on IMO 2025 and CMO 2024, and 118/120 on Putnam 2024 with scaled test-time compute.
The larger lesson is more important than the competition scores:
A self-improving generator eventually creates examples that its old verifier was never trained to judge. Verification has to improve with the generator.
4. The verifier can become the new reward-hacking surface
This creates a recursive problem.
If a system is optimized against a verifier, the generator can eventually learn to exploit weaknesses in that verifier.
A stronger generator therefore does not automatically mean a stronger self-improvement loop.
The system also needs ways to audit the verifier itself.
This is why the lecture’s emphasis on meta-verification matters.
In a mature self-improvement system, the questions become:
- Is the candidate solution correct?
- Is the verifier’s criticism correct?
- Is the reward signal aligned with the task we actually care about?
- Has the model learned the task, or merely learned the evaluator?
This is a recurring pattern across the entire series.
The bottleneck moves from generation to selection, then from selection to the reliability of the selector.
5. Bottleneck three: what happens when humans can no longer provide the best next task?
Even a strong generator and verifier still need useful problems to train on.
Today, many reinforcement-learning pipelines depend on human-curated datasets or tasks selected in advance.
That creates a scaling limit.
If the model becomes stronger than the available task distribution, additional training can become repetitive or too easy.
The mechanism described in the lecture closely matches Absolute Zero: Reinforced Self-play Reasoning with Zero Data (arXiv:2505.03335).
Absolute Zero uses one model in two roles:
- Proposer: generate new reasoning tasks.
- Solver: solve those tasks.
The environment validates tasks and answers with executable code.
The system uses three reasoning modes:
- Deduction: infer output from program and input.
- Abduction: infer a valid input from program and output.
- Induction: infer a program from input-output examples.
The proposer is rewarded for learnability: tasks should not be trivial, but they should still be solvable enough to provide useful learning signal.
This creates an automatic curriculum that changes as the solver changes.
6. Why self-generated curriculum is more important than “zero data”
The phrase “zero data” can be misleading if taken too literally.
Absolute Zero still starts from a pretrained model and depends on an external execution environment to determine whether generated tasks and answers are valid.
Its contribution is not that learning occurs without any prior knowledge or external structure.
The important change is where the next training task comes from.
Instead of selecting from a fixed human-curated problem set, the model helps construct the next edge of its own curriculum.
That changes the role of the environment.
The environment becomes both:
- a source of objective feedback,
- and a constraint that keeps open-ended task generation grounded.
This is powerful in code and mathematics because execution can provide relatively hard verification.
It is much harder in open-ended domains where there is no cheap deterministic checker.
7. The strongest counterargument: self-play works best where reality is easy to verify
This is the most important limitation of the self-improvement story.
Code can be executed.
Many mathematical results can be checked.
Games have explicit rules.
But many economically valuable tasks do not have a cheap ground-truth verifier.
Consider:
- strategy,
- management,
- scientific hypothesis generation,
- policy analysis,
- product design,
- negotiation,
- medical judgment,
- long-term organizational decisions.
In these domains, the quality of an answer may only become clear after days, months, or years.
That means the self-improvement loop faces an asymmetry:
generation can become cheap faster than verification becomes cheap.
The system can create enormous quantities of candidate reasoning long before it can reliably determine which candidates should become future training data.
So the verifier bottleneck is not a temporary implementation detail.
It may be one of the fundamental limits of open-ended self-improvement.
8. A fourth bottleneck: intelligence still has to run on physical hardware
The second half of the lecture turns from learning loops to systems efficiency.
This matters because smarter agents can consume more inference-time compute, more memory, and more energy.
A research direction called Intelligence per Watt asks a practical systems question:
How much useful task performance do we obtain per unit of power?
The verified paper is arXiv:2511.07885, “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI.”
It evaluates:
- 20+ local language models
- 8 hardware accelerators
- 1 million real-world single-turn chat and reasoning queries
The paper defines intelligence per watt as task accuracy divided by power.
Its headline results are more nuanced than “move everything to local devices.”
9. Local AI is improving rapidly, but cloud accelerators are still more efficient on the same workload
The paper reports that local models could correctly answer 88.7% of the sampled single-turn queries when considering the available local-model set.
From 2023 to 2025:
- local query coverage increased from 23.2% to 71.3%
- intelligence per watt improved 5.3×
The decomposition is approximately:
- 3.1× from better models
- 1.7× from better accelerators
That is different from saying hardware and algorithms contributed exactly 50/50.
The paper also finds that local accelerators still have at least 1.4× lower IPW than cloud accelerators when running the same models.
So the result is not “laptops beat datacenters.”
The more interesting result is architectural.
A hybrid system can keep easier queries local and escalate harder queries to the cloud.
The gain comes from routing the right task to the right model and hardware, not from choosing one substrate for everything.
10. Self-improvement and efficiency are the same systems problem at different layers
At first glance, multiagent training and energy-efficient inference look like separate research topics.
They are connected by resource allocation.
A self-improving system constantly makes allocation decisions:
- Which candidate reasoning paths deserve more compute?
- Which verifier should inspect them?
- Which generated tasks are worth training on?
- Which model should handle a user request?
- Which hardware should execute it?
- When should a local model escalate to a frontier model?
The common problem is not simply intelligence.
It is selective expenditure of scarce resources.
Compute, verification effort, training data, memory, energy, and human attention all have opportunity costs.
A mature self-improving agent has to learn not only how to reason, but where reasoning effort is worth spending.
11. Continual learning remains an unresolved bridge
The lecture closes with another difficult question.
Can external memory and retrieval replace continual weight updates?
For factual recall, retrieval can often be enough.
But the lecture argues that new reasoning patterns, domain adaptation, and cross-embodiment skill transfer may require updates to the underlying model or policy.
This claim should be treated as a research direction rather than a settled universal law.
External memory, in-context adaptation, tool use, parameter-efficient updates, and full continual learning occupy different points in the design space.
The open problem is to determine what should be stored outside the model and what should change the model itself—without catastrophic forgetting or destabilizing previously learned capabilities.
12. What this article does not claim
- It does not claim that multiagent fine-tuning guarantees unlimited self-improvement.
- It does not claim that a stronger verifier eliminates reward hacking.
- It does not claim that self-generated tasks remove the need for pretrained knowledge or external environments.
- It does not claim that code-based self-play directly transfers to open-ended real-world domains.
- It does not claim that local hardware is more efficient than datacenter accelerators on identical workloads.
- It does not claim that external memory and continual weight updates are mutually exclusive.
- It does not claim that this nine-video archive represents the complete official Stanford CS329A course schedule.
13. The real thesis of Part 9
The most important shift in the series is from model-centric improvement to loop-centric improvement.
A useful self-improving system needs at least four coupled capabilities:
1. Generate diverse possibilities
Avoid collapsing into a narrow family of reasoning patterns.
2. Verify increasingly difficult outputs
Prevent the generator from outrunning the evaluator.
3. Generate the next useful training problem
Keep the curriculum near the frontier of the system’s current ability.
4. Allocate compute and energy intelligently
Use the right model, verifier, and hardware for each task.
These loops depend on each other.
More generation without verification produces more noise.
More verification without new tasks produces stagnation.
More tasks without diversity produces repetition.
More capability without efficiency makes the system difficult to deploy.
That leads to the final thesis of this nine-video series:
The path to self-improving agents is not a single model teaching itself forever. It is a coordinated system that learns what to generate, what to trust, what to practice next, and where computation is worth spending.
Claim Map
Lecture claims
- diversity collapse can limit repeated single-model self-training
- verification becomes harder as generators become stronger
- self-generated curricula can reduce dependence on human task curation
- continual learning and efficient inference remain major open research areas
Verified facts used in this article
- Multiagent Finetuning: arXiv:2501.05707
- DeepSeekMath-V2: arXiv:2511.22570
- Absolute Zero: arXiv:2505.03335
- Intelligence per Watt: arXiv:2511.07885
- Intelligence per Watt evaluates 20+ local LMs, 8 accelerators, and 1M real-world queries
- local-model set covers 88.7% of sampled single-turn queries
- IPW improved 5.3× from 2023 to 2025, decomposed into approximately 3.1× model and 1.7× accelerator gains
- local accelerators remain at least 1.4× lower in IPW than cloud accelerators on identical-model comparisons
Editorial interpretation
- future self-improvement is fundamentally a coordination problem across generation, verification, curriculum, and compute allocation
- delegating more reasoning to AI increases the importance of verification economics rather than reducing it
Lecture Map
- 00:00–00:15 — self-improvement bottlenecks and multiagent diversity
- 00:15–00:23 — verification and self-verifiable mathematical reasoning
- 00:23–00:34 — autonomous task proposal and solver loops
- 00:40–00:52 — intelligence per watt and local/cloud inference
- 00:52–01:08 — continual learning, memory, and closing research questions
Primary Sources
- Stanford CS329A Part 9 video: https://www.youtube.com/watch?v=z3q9aQUsgRs
- Stanford CS329A: https://cs329a.stanford.edu/
- Multiagent Finetuning: https://arxiv.org/abs/2501.05707
- DeepSeekMath-V2: https://arxiv.org/abs/2511.22570
- Absolute Zero: https://arxiv.org/abs/2505.03335
- Intelligence per Watt: https://arxiv.org/abs/2511.07885