Series: Self-Improving AI Agents — Part 7 of 9
Previous: Part 6 — Train-Time Scaling
Lecture: Stanford CS329A | Self-Improvement and Deep Research Agents
Source video: https://www.youtube.com/watch?v=Uni9dqyuuDM
1. Part 7 moves from solving to searching
Part 6 focused on changing model behavior through training. Part 7 examines another route to improvement: keep the base model fixed, but search more intelligently over possible outputs and external knowledge.
The lecture develops two forms of search.
- Output-space search: generate many candidate programs, test them, cluster them, and select a small set of promising solutions. AlphaCode and AlphaCode 2 are the main examples.
- Knowledge-space search: let a reasoning model retrieve external information during reasoning, refine what it found, and inject only useful evidence back into the reasoning chain. Search-o1 and Search-R1 are the main examples.
The common pattern is:
generate or retrieve many possibilities → filter noise → rank or refine candidates → continue with a smaller, better set
The central thesis is:
Search only becomes intelligence when the system can decide what to keep, what to discard, and how to integrate the surviving evidence.
2. AlphaCode: massive sampling plus behavioral selection
Competitive programming is much harder than short function completion. A problem can require long natural-language specifications, hidden constraints, algorithm selection, data-structure design, and a complete executable program.
AlphaCode attacked this by generating an enormous number of candidate programs and then aggressively reducing them.
Its system combined code pretraining, competitive-programming fine-tuning, large-scale C++ and Python sampling, execution on public tests, behavioral clustering, reranking, and submission of only a small final set.
In simulated Codeforces contests, AlphaCode achieved an average estimated ranking in the top 54.3% of participants.
The important point is not simply “one million samples.” The important point is that massive sampling was paired with a selection mechanism based on program behavior.
3. Pass@K versus limited submissions
If a correct program appears somewhere among many candidates, Pass@K can be high. But a real contest does not allow unlimited submissions.
That creates two different problems:
- Coverage: did the generator produce a correct answer somewhere?
- Selection: did the system identify that answer among the candidates?
This distinction generalizes beyond coding. Many agents can occasionally produce a good solution. The operational challenge is making the good solution surface reliably before compute, latency, or user patience runs out.
4. AlphaCode 2 improved both proposal quality and ranking
AlphaCode 2 changed the system in several ways.
It used Gemini Pro as the foundation for its components, sampled only C++, used a family of fine-tuned policy models to preserve diversity, and added a learned scoring model.
The scoring model estimates whether a candidate program is likely to be correct. This gives the system a learned reranker rather than relying only on clustering heuristics.
The official report says AlphaCode 2:
- solved 43% of the evaluated problems versus 25% for AlphaCode,
- reached about the 85th percentile on Codeforces,
- and needed about 100 samples to reach the performance level that AlphaCode required one million samples to achieve.
The deeper lesson is:
Better search requires both a better proposal distribution and a better evaluator.
5. The 95% filtering figure is not a compile-failure rate
The lecture-related summary compresses the AlphaCode 2 filtering stage too aggressively.
The technical report says less than 5% of samples fail to compile. The broader filtering stage removes about 95% of candidates because they either fail the public input/output tests or otherwise cannot be correct.
So the lesson is not that Gemini-based generation produced almost entirely uncompilable code. It is that most candidate programs were behaviorally wrong and had to be eliminated before clustering and scoring.
6. Search has diminishing returns when diversity saturates
Sampling more candidates helps only while additional samples produce meaningfully different solution strategies.
Eventually a weak base model may generate variants of the same wrong idea. At that point, more search budget increases cost faster than useful coverage.
This is one reason AlphaCode 2’s stronger foundation model matters: it increases the density of useful candidates before the filtering stage.
Search cannot recover a solution that the generator never makes reachable.
7. Deep research agents apply the same logic to external knowledge
The second half of Part 7 moves from program search to information search.
Large reasoning models can reason over long chains, but they still encounter knowledge gaps. Standard RAG retrieves documents before generation and places them into context. For complex reasoning, that can be insufficient.
A model may discover a new information need halfway through reasoning. Agentic RAG allows the model to search dynamically, but that introduces another problem: context pollution.
If every retrieved page is appended in full, context grows while the useful signal can become harder to isolate.
8. Search-o1 adds a separate Reason-in-Documents stage
Search-o1 combines agentic retrieval with a separate Reason-in-Documents module.
The reasoning model can decide when it needs external knowledge and issue a search query. Retrieved pages are not simply dumped into the main reasoning chain.
A separate refinement stage analyzes:
- the current search query,
- the retrieved documents,
- and the reasoning accumulated so far.
It then extracts concise information relevant to the current reasoning need and injects that refined result back into the main chain.
This is the key design choice: raw retrieval is treated as intermediate material, not finished context.
9. More documents help only when the system can refine them
The Search-o1 paper reports that Search-o1 can benefit from increasing the number of retrieved documents on complex reasoning tasks, while simpler retrieval approaches do not benefit in the same way.
More retrieval increases both useful evidence and noise. Without refinement, both grow together.
So the correct lesson is not “longer context is better.”
The value of retrieval depends on whether the system can compress raw evidence into the exact form the next reasoning step needs.
10. Human-expert comparisons are domain-specific
The lecture-related summary emphasizes strong GPQA performance, but the public claim needs precision.
On the GPQA extended set reported in Search-o1:
- Physics: Search-o1 68.7, physicist comparison 57.9
- Biology: Search-o1 69.5, biologist comparison 68.9
- Chemistry: Search-o1 40.7, chemist comparison 72.6
So Search-o1 does not establish that the system “beats human experts in science.” It shows strong, domain-dependent performance: above the reported human comparison in Physics and Biology, but far below in Chemistry.
11. Search-o1 to Search-R1: from scaffolded search to learned search policy
Search-o1 is primarily an inference-time orchestration framework.
Search-R1 moves the search behavior into reinforcement learning.
The recorded source metadata lists the wrong identifier. The verified paper is arXiv:2503.09516, not 2502.12345.
Search-R1 trains models to generate multiple search queries during step-by-step reasoning and interact with retrieval through RL. The paper reports improvements over prior baselines across seven QA datasets, including 26% for Qwen2.5-7B, 21% for Qwen2.5-3B, and 10% for LLaMA3.2-3B in its reported setup.
The conceptual transition is:
- Search-o1: a system designer specifies the search-and-refine loop.
- Search-R1: the model learns a search policy from reward.
This reconnects Part 7 with Part 6: search begins as scaffolding and becomes trainable behavior.
12. The strongest counterargument: is this self-improvement, or just a better wrapper?
AlphaCode, AlphaCode 2, and Search-o1 improve end-to-end performance partly through components outside the base language model:
- sampling budget,
- execution environments,
- clustering,
- scoring models,
- retrieval engines,
- search APIs,
- context-refinement modules.
A benchmark gain therefore cannot automatically be attributed to improved internal reasoning in the base model.
A stronger scaffold can make the same model much more capable as a system.
That is not a weakness if the goal is useful agents. But it changes the claim.
Agent capability is a property of the model-plus-search-plus-evaluation system, not only of the model weights.
13. Search quality is bounded by proposal quality and evaluator quality
Part 7 can be reduced to two bottlenecks.
Proposal bottleneck
Can the system generate or retrieve a candidate containing the right solution or evidence?
Evaluation bottleneck
Can the system recognize that candidate as better than the alternatives?
AlphaCode’s massive sampling attacks the first bottleneck. AlphaCode 2 improves both. Search-o1 refines raw documents before they enter the reasoning chain. Search-R1 learns a policy for when and how to search.
If either side fails, more compute can create more noise instead of more intelligence.
14. What this article does not claim
- The lecture did not show that one million samples is a practical strategy for every coding assistant.
- It did not show that 95% of AlphaCode 2 samples fail to compile.
- It did not show that Search-o1 outperforms human experts in every science domain.
- It did not show that retrieval volume alone improves reasoning.
- It did not show that a scaffolded agent’s benchmark gain is identical to an improvement in base-model intelligence.
- It did not establish that Search-R1 is the final form of autonomous research agents.
15. The real thesis of Part 7
Part 6 showed how successful trajectories can be converted into training signal.
Part 7 shows where those trajectories can come from: search.
But search is useful only when the system can reduce the search space again.
In AlphaCode, execution and clustering reduce many programs to a few submissions.
In AlphaCode 2, a stronger foundation model and learned scorer make that compression much more efficient.
In Search-o1, raw documents are refined into evidence that the current reasoning step can use.
In Search-R1, the search policy itself becomes trainable.
The deeper pattern is:
Self-improving agents need two loops: a loop that explores possibilities, and a loop that learns which possibilities deserve to survive.
Part 8 moves to Agentic Evaluations and Long-Horizon Tasks. Unlike Part 5’s broad distinction between capability and reliability, Part 8 asks how task duration, reliability, professional output quality, and verifiability interact at deployment scale.
Claim Map
Speaker / course claims
- Large-scale search can substantially improve coding and research-agent performance.
- Selection and context refinement are critical bottlenecks.
- Search behavior can move from prompt-engineered scaffolding toward RL-trained policies.
Verified facts
- AlphaCode is arXiv:2203.07814 and reached an estimated top-54.3% ranking in simulated Codeforces contests.
- AlphaCode 2 solved 43% of evaluated problems versus 25% for AlphaCode and reached about the 85th percentile.
- AlphaCode 2 requires about 100 samples to reach AlphaCode’s performance level at one million samples.
- AlphaCode 2 filtering removes about 95% of candidates overall, while fewer than 5% fail to compile.
- Search-o1 is arXiv:2501.05366 and uses agentic retrieval plus Reason-in-Documents.
- Search-o1’s GPQA expert comparison is strong in Physics and Biology but weak in Chemistry.
- Search-R1 is arXiv:2503.09516 and trains multi-turn search behavior with reinforcement learning.
Editorial interpretation
- Search is useful only when proposal quality and evaluator quality improve together.
- System-level agent capability should be distinguished from base-model capability.
Lecture Map
- 00:00–00:25 — AlphaCode and output-space search
- 00:25–00:46 — AlphaCode 2, diversity, filtering, and learned scoring
- 00:46–01:07 — Search-o1, agentic RAG, and Reason-in-Documents
- 01:07–01:12 — Search-R1 and learned search policies
Primary Sources
- Stanford CS329A Part 7 video: https://www.youtube.com/watch?v=Uni9dqyuuDM
- Stanford CS329A: https://cs329a.stanford.edu/
- AlphaCode: https://arxiv.org/abs/2203.07814
- AlphaCode 2 Technical Report: https://deepmind.google/AlphaCode2_Tech_Report.pdf
- Search-o1: https://arxiv.org/abs/2501.05366
- Search-R1: https://arxiv.org/abs/2503.09516