Series: Self-Improving AI Agents — Part 3 of 9
Previous: Part 2 — Test-Time Compute Is a Resource Allocation Problem
Lecture: Stanford CS329A | Robust Verification
Part 2 ended with a constraint that test-time scaling cannot escape: generating more candidate solutions is useful only if the system can identify which candidates are actually correct.
Part 3 turns that constraint into the main subject.
The lecture traces a four-stage evolution in verification research: outcome verifiers that score completed solutions, process reward models that judge intermediate steps, automatically generated process supervision, and weak-verifier ensembles that combine many imperfect judges.
The common problem underneath all four is simple to state and hard to solve:
What happens when the verifier—the component deciding what is correct—is itself fallible?
This question matters because verification is not just a filter at the end of inference. Once verifier scores are used to guide search, reinforcement learning, or self-improvement, verifier errors can become training signals. A weak verifier can therefore do more than select the wrong answer: it can teach the system to become better at producing answers that look correct to the verifier.
1. The first step: train a model to rank solutions, not just generate them
The lecture begins with OpenAI’s 2021 paper Training Verifiers to Solve Math Word Problems, which introduced GSM8K, a dataset of 8.5K grade-school math word problems designed to require multi-step reasoning.
The central idea was to separate generation from verification.
- Fine-tune a generator on worked solutions.
- Sample many candidate completions for each problem.
- Label those completions by whether they reach the correct final answer.
- Train a separate verifier to score candidate solutions.
- At test time, generate multiple solutions and return the one ranked highest by the verifier.
The paper reports that verification scales more effectively with additional data than a fine-tuning-only baseline. In the original experiments, a 6B model using verification reached performance comparable to a much larger 175B fine-tuned generator in the evaluated GSM8K setting.
This was an early demonstration of a recurring theme in the series: more parameters in the generator are not the only way to buy better task performance.
2. But outcome verification contains a hidden loophole
The 2021 verifier is trained primarily from the final answer.
If a solution reaches the correct number, it can receive a positive label even if the reasoning contains a flawed step. The authors explicitly note that final-answer labeling introduces false positives when a completion arrives at the right result through incorrect reasoning.
This creates a fundamental ambiguity.
Did the model solve the problem correctly, or did it make an error and later cancel that error by accident?
For simple reranking, that distinction may sometimes be tolerable. For self-improvement, it is dangerous. If the verifier rewards a flawed trajectory because its final answer is correct, reinforcement learning can increase the probability of that flawed reasoning pattern.
The key lesson is:
a verifier that only sees the destination can miss corruption in the path.
3. More search can eventually exploit verifier weakness
The lecture also emphasizes an uncomfortable property of large candidate pools.
Suppose the verifier is very good but not perfect. With a small number of samples, the highest-scoring candidate is often genuinely strong. But as the candidate pool grows, the system gets more opportunities to produce an answer that is wrong in a subtle way yet happens to receive an unusually high verifier score.
This is a form of selection pressure against the verifier.
The more aggressively we optimize for the verifier’s score, the more likely we are to discover its blind spots.
The original GSM8K verifier experiments show that performance does not increase indefinitely with the number of candidate completions; the lecture highlights a peak-and-decline regime at large sample counts. The exact optimum depends on the setup, but the general failure mode is more important than the number:
search can become an adversary of an imperfect evaluator.
This is the same structural problem that later appears under names such as reward hacking, specification gaming, and evaluator exploitation.
4. Process supervision changes what gets verified
The next major step is OpenAI’s 2023 paper Let’s Verify Step by Step.
Instead of assigning feedback only to a final answer, process supervision provides feedback to intermediate reasoning steps. The authors trained process reward models on the MATH dataset and found that process supervision significantly outperformed outcome supervision in that setting.
The paper also released PRM800K, a dataset of 800,000 step-level human feedback labels used to train the best process reward model.
The conceptual shift is important:
- Outcome supervision: Was the final answer correct?
- Process supervision: Was each intermediate step justified?
If an error appears halfway through a solution, a process verifier can penalize that trajectory before the final answer is produced.
That makes process supervision useful not only for ranking finished responses but also for tree search, beam search, revision, and reinforcement learning.
5. Process supervision is stronger—but not free
Step-level labels are expensive.
A single solution can contain many reasoning steps, and each step may require a human judgment. The verification signal becomes more precise, but the annotation cost rises sharply.
The Lightman et al. study therefore explores active learning: instead of labeling random examples uniformly, prioritize examples that are likely to be informative, especially convincing wrong solutions.
In a controlled synthetic experiment, the paper estimates that active learning improves data efficiency by roughly 2.6× over uniform sampling. That figure should be interpreted narrowly: it is not a universal law about human labeling costs, but evidence that targeted selection of difficult examples can make process supervision more efficient.
This creates the next research question:
Can we create process supervision without paying humans to label every step?
6. Math-Shepherd tries to automate the process labels
Math-Shepherd, published at ACL 2024, attacks exactly that bottleneck.
The method assigns rewards to intermediate mathematical reasoning steps using automatically constructed process-supervision data rather than relying entirely on human step labels.
The core idea is rollout-based evaluation. Given an intermediate step, continue the solution multiple times. If continuations from that step frequently reach the correct final answer, the step is treated as more promising.
This is not a perfect definition of correctness. A bad step may still be recoverable, and a good step may lead to failure because of later mistakes. But the rollout provides a measurable proxy for the value of an intermediate state.
Math-Shepherd uses this signal in two ways:
- Verification: rerank multiple generated solutions.
- Reinforcement learning: use step-level rewards to improve the generator.
The paper reports that process RL with Math-Shepherd improved Mistral-7B from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on MATH. Adding Math-Shepherd verification further raised the reported results to 89.1% and 43.5%, respectively.
The significance is not that human supervision becomes unnecessary everywhere. The stronger conclusion is that some domains with objective final answers can bootstrap intermediate feedback from their own rollout structure.
7. Automatic process supervision still depends on ground truth
That qualification matters.
Math-Shepherd can automate step labels because the final mathematical answer can be checked. Rollouts have a destination against which success can be measured.
In open-ended domains—strategy, policy, scientific interpretation, creative work, or ambiguous real-world decisions—there may be no cheap ground-truth endpoint.
This means that automatic process supervision does not eliminate the verification problem. It moves the burden.
Instead of asking humans to label every step, the system relies on a trustworthy final-outcome signal.
Where that signal does not exist, the system needs another source of judgment.
8. The lecture’s final move: combine imperfect verifiers instead of waiting for a perfect one
The last research line in the lecture is the Stanford-led 2025 paper Shrinking the Generation-Verification Gap with Weak Verifiers.
The primary paper calls the framework Weaver.
Its premise is pragmatic: in many realistic tasks, we do not have a perfect verifier. What we do have is a collection of imperfect reward models, LM judges, and other scoring functions. If their errors are not identical, their combined signal may be stronger than any single judge.
Weaver therefore:
- collects scores from multiple weak verifiers,
- normalizes inconsistent score formats,
- filters low-quality verifiers,
- uses weak supervision to estimate relative verifier reliability, and
- combines the outputs into a stronger aggregate score.
The critical idea is not simple majority voting. A weak verifier that is systematically better should receive more weight than a noisy one.
9. Weak supervision can recover signal without large labeled verifier datasets
Weighted verifier ensembles would normally require labeled examples to learn which judges deserve trust.
Weaver tries to reduce that dependency using weak supervision. Rather than requiring a large ground-truth dataset, it estimates verifier quality from patterns of agreement, disagreement, and dataset statistics.
This is powerful, but it introduces an assumption: the verifier errors must contain enough independent information to be disentangled.
If every verifier is built from the same model family, trained on the same data, and fails in the same way, an ensemble can become highly confident and still be wrong.
So diversity among verifiers is not automatically useful. What matters is diversity of error structure.
10. The verified Weaver result is strong—and different from the source metadata recorded for this lecture
The primary source matters here because the lecture archive contains a naming and metadata mismatch.
The paper’s framework is Weaver, not “Beaver,” and its arXiv identifier is 2506.18203.
In the paper’s reported evaluation, Weaver paired with Llama 3.3 70B Instruct as the generator and an ensemble of verifiers no larger than 70B achieved 87.7% average accuracy across the evaluated reasoning and math tasks—described by the authors as o3-mini-level performance in that benchmark mix.
The point is not that a weak-verifier ensemble universally “matches o3-mini.” The result is benchmark- and setup-specific. But it demonstrates that verifier quality can recover a large fraction of the gap between first-sample performance and the generator’s latent coverage.
11. Verifier ensembles introduce their own compute problem
Calling dozens of large verifier models for every candidate solution is expensive.
Weaver therefore adds a second step: distillation.
The full verifier ensemble produces combined scores, and a much smaller 400M cross-encoder is trained to imitate those scores.
Different revisions and figures of the paper report slightly different retention summaries, so the safest characterization is approximate: the distilled verifier preserves about 98% of the ensemble’s performance while cutting verification compute by as much as 99.97% in the reported setup.
This is a significant systems insight.
Verification can be expensive during development and cheap at deployment if the aggregate judgment of many teachers can be compressed into one small student verifier.
12. But distillation can also compress the ensemble’s blind spots
Compression does not create truth.
If the teacher ensemble contains a shared bias, the student verifier can faithfully learn that bias. Distillation improves efficiency; it does not guarantee correctness.
This is a recurring pattern throughout the lecture:
- Outcome verifiers can reward flawed reasoning.
- Process verifiers can be expensive and reward-hacked.
- Rollout-based supervision depends on a reliable final answer signal.
- Weak-verifier ensembles depend on sufficiently independent errors.
- Distilled verifiers inherit the strengths and weaknesses of their teachers.
Robust verification is therefore not a single algorithm. It is a chain of assumptions that must remain valid under stronger search pressure.
13. The strongest counterargument: verification may simply move the alignment problem one layer down
A skeptical reader can ask a deeper question.
If we cannot trust the generator, we build a verifier. If we cannot trust the verifier, we build a process verifier, an ensemble, or a meta-verifier. But what verifies the verifier?
There is no universal answer.
In formal mathematics, code execution, or theorem proving, external tools can provide relatively hard signals. In open-ended reasoning, evaluator quality becomes part of the system’s epistemic limit.
This means that “more verification” is not automatically equivalent to “more truth.” It can also mean deeper dependence on a scoring function that the optimizer may eventually exploit.
The defensible conclusion is narrower:
verification improves self-improving systems when the feedback channel is harder to game than the behavior it is supervising.
That is the real standard.
14. What this article does not claim
- The lecture did not claim that process reward models completely replace outcome reward models in every setting.
- It did not claim that automated process supervision works without any trustworthy outcome signal.
- It did not establish that an ensemble of weak verifiers is reliable when all verifiers share the same failure mode.
- It did not show that math verification results transfer unchanged to subjective domains.
- It did not imply that a smaller distilled verifier becomes more correct than the teacher ensemble simply because it is cheaper.
- It did not eliminate reward hacking; stronger optimization can expose weaknesses in the verifier itself.
15. The real thesis of Part 3
The history covered in this lecture can be read as a sequence of attempts to improve the feedback channel:
- Outcome verifier: rank finished solutions.
- Process verifier: inspect intermediate reasoning.
- Automatic process supervision: derive step-level signals from rollouts.
- Weak-verifier ensemble: combine multiple imperfect judges.
- Distillation: compress expensive verification into a deployable model.
Each step makes verification more scalable. None makes it infallible.
That is why robust verification is central to self-improving agents. Once a verifier controls search or learning, its errors stop being passive mistakes. They become directions for optimization.
The core question is therefore not “Can the system score its own answers?” It is:
Can the feedback mechanism remain trustworthy as the generator becomes increasingly good at optimizing against it?
Part 4 will move from judging outputs to interacting with the world: how can tools, code execution, and environment feedback provide stronger learning signals than model-based judgment alone?
Claim map
- Speaker / course claims: verification is the central bottleneck exposed by repeated sampling; process supervision can reduce false-positive reasoning errors; automatic process supervision and weak-verifier ensembles are paths toward scalable verification.
- Verified facts: GSM8K contains 8.5K math word problems; Cobbe et al. train verifiers to rank many candidate solutions and show verification scales better with data than fine-tuning alone; Lightman et al. find process supervision outperforms outcome supervision on the cited MATH setting and release PRM800K with 800,000 step-level labels; Math-Shepherd automatically constructs process supervision and reports the stated Mistral-7B improvements; Weaver combines weak verifiers, reports 87.7% average accuracy in its evaluated setup with a Llama 3.3 70B generator, and distills its ensemble into a 400M cross-encoder.
- Editorial interpretation: “verification is a feedback channel that must remain harder to game than the behavior it supervises” is the organizing synthesis of this article, not a direct quotation from one paper.
Lecture map
- 00:00–00:21 — GSM8K and outcome verifiers
- 00:21–00:38 — process supervision and PRM800K
- 00:38–00:52 — Math-Shepherd and automatic step supervision
- 00:52–01:07 — weak-verifier ensembles and distillation
- 01:07–01:13 — verifier limits and future directions
Primary sources
- Stanford CS329A — Part 3 | Robust Verification
- Stanford CS329A — Self-Improving AI Agents
- Training Verifiers to Solve Math Word Problems
- Let’s Verify Step by Step
- Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
- Shrinking the Generation-Verification Gap with Weak Verifiers