Train-Time Scaling: How AI Learns from Its Own Reasoning — Stanford CS329A Part 6

Series: Self-Improving AI Agents — Part 6 of 9

Previous: Part 5 — Capability Is Not Reliability

Lecture: Stanford CS329A | Train-Time Scaling

Source video: https://www.youtube.com/watch?v=7xr620o0sGM

Series note: “Part 6” follows the nine-video archive used for this series.

한국어판

1. Test-time scaling asks for more thinking. Train-time scaling changes the model.

Earlier parts of this series focused on what an agent can do at inference time: sample multiple answers, use tools, verify outputs, and spend more compute on hard problems.

Train-time scaling closes a different loop.

Instead of throwing away the successful trajectories after inference, the system uses them as training signal. The model generates candidate reasoning, a verifier decides which outputs are useful, and those outputs are fed back into the next training round.

The basic cycle is:

generate → verify → select → update weights → generate again

That makes train-time scaling more than “reasoning longer.” It is an attempt to turn successful reasoning into a persistent change in the model’s output distribution.

The lecture develops this idea through three research lines:

  1. STaR — bootstrap reasoning traces from the model’s own successful outputs.
  2. DeepSeekMath / GRPO — use group-relative rewards to improve mathematical reasoning without a separate critic model.
  3. DAPO — stabilize large-scale RL when long reasoning causes exploration collapse, zero-gradient batches, and length pathologies.

The central thesis is:

Self-improvement becomes durable when successful reasoning is not only selected at inference time, but converted into training data or reinforcement-learning updates.

2. STaR begins with a simple problem: good rationales are expensive

Reasoning models benefit from step-by-step rationales, but high-quality rationale datasets are expensive to create.

STaR — Self-Taught Reasoner — proposes a bootstrap loop.

A model receives a small number of rationale examples. It then generates rationales for a larger set of problems. When the final answer is correct, that generated rationale can be treated as training data.

The model is fine-tuned on those successful traces and then asked to solve the dataset again.

The point is not that every generated rationale is trustworthy. The point is that a model can use its own successful outputs to enlarge the set of reasoning traces available for training.

STaR’s loop is therefore:

few rationale examples → generate reasoning → keep successful traces → fine-tune → repeat

The original paper reports substantial improvements over directly fine-tuning on final answers alone and shows that the approach can work on both mathematical and commonsense reasoning tasks.

3. Rationalization with hints expands the training set — but creates a new risk

STaR adds a clever mechanism for problems the model initially fails.

If the model does not reach the correct answer, the correct answer is supplied as a hint and the model is asked to generate a rationale that explains how that answer could be reached.

If the resulting rationale leads to the right answer, the hint is removed and the rationale can be added to the training set.

This is powerful because failed examples no longer have to be discarded immediately.

But it also creates a structural danger.

A model that already knows the answer can produce a persuasive post hoc explanation that is not the reasoning path it would have discovered independently.

That means “final answer correct” is not equivalent to “reasoning trace correct.”

This is one of the first places where Part 6 connects back to Part 3: the quality of self-improvement depends on the quality of the verifier.

4. STaR does not prove that a model can bootstrap arbitrary new knowledge

The strongest interpretation of STaR would be that a model can recursively teach itself anything.

The evidence does not justify that claim.

STaR is better understood as a method for extracting and reinforcing reasoning patterns that are already reachable from the base model plus the provided supervision.

The method still depends on:

  • an initial capable model,
  • problems with known answers,
  • some mechanism for deciding whether outputs are correct,
  • and a distribution where successful rationales can actually be generated.

This becomes important later when we examine DeepSeekMath.

5. DeepSeekMath shows that the base model and data still matter before RL

The lecture materials correctly identify DeepSeekMath as a major transition point, but one numerical detail requires correction.

The 51.7% result reported by DeepSeekMath is on the MATH benchmark, not AIME.

The verified DeepSeekMath paper reports that DeepSeekMath 7B continues pre-training from DeepSeek-Coder-Base-v1.5 7B using 120 billion math-related tokens from Common Crawl, together with natural-language and code data.

This matters because the later RL results do not start from an empty model.

The base model has already been shaped by code pretraining, mathematical web data, and instruction tuning.

Only after that does GRPO change the output distribution.

So the train-time story is not:

RL replaces pretraining.

It is closer to:

pretraining creates reachable capabilities; post-training changes how reliably the model selects and expresses them.

6. GRPO replaces the critic with relative comparison inside a group

PPO-style reinforcement learning commonly uses a learned value model — a critic — to estimate expected return and calculate an advantage.

For a large language model, another model of similar size can become a substantial memory and compute burden.

GRPO — Group Relative Policy Optimization — changes the baseline.

For each prompt, the policy generates a group of outputs. Those outputs receive rewards. Instead of asking a separate critic how good each response was expected to be, GRPO compares each response with the reward distribution of its siblings.

A simplified group-relative advantage is:

advantage ≈ (reward − group mean) / group standard deviation

An answer that performs better than the group gets reinforced; one below the group is discouraged.

The key architectural shift is that the group becomes the comparison baseline, so a separately trained value model is no longer required.

7. DeepSeekMath improved top-1 performance — but the Pass@K result is more revealing

DeepSeekMath reports a concrete improvement from its RL stage.

On MATH, the instruction model improves from roughly 46.8% to 51.7% after GRPO-based RL. With self-consistency over 64 samples, the paper reports 60.9%.

Those are important results.

But the paper’s deeper experiment compares Maj@K and Pass@K.

  • Maj@K asks whether the majority answer among K samples is correct.
  • Pass@K asks whether at least one of the K samples is correct.

The paper finds that RL improves Maj@K much more clearly than Pass@K as K grows.

The authors interpret this as evidence that, in their setup, RL mainly makes correct responses more likely among answers the model could already produce, rather than dramatically expanding the set of problems for which a correct solution is reachable.

That distinction matters.

RL can make a model much more useful without proving that it created entirely new underlying knowledge.

Moving probability mass toward good trajectories is operationally valuable even when the frontier of reachable solutions changes little.

8. This creates a more precise definition of train-time scaling

The phrase “train-time scaling” can sound like simply spending more GPUs on training.

Part 6 points to something more specific.

The important resource is not only compute. It is the closed-loop production of informative training signal.

A useful loop needs:

  1. enough exploration to generate different candidate trajectories,
  2. a reward or verifier that can distinguish better outputs,
  3. an update rule that uses those differences efficiently,
  4. and enough stability to avoid collapsing the policy.

That final requirement is what DAPO addresses.

9. The DAPO source metadata recorded for this lecture is wrong

The source metadata recorded for this lecture lists DAPO as arXiv 2501.12345.

The verified paper is:

DAPO: An Open-Source LLM Reinforcement Learning System at Scale — arXiv:2503.14476

The paper was submitted in March 2025 and later appeared at NeurIPS 2025.

This correction is material because Part 6 depends heavily on DAPO’s exact training design and ablation results.

The public article should therefore use 2503.14476.

10. DAPO shows why large-scale reasoning RL is a systems problem

Naive GRPO is not automatically stable when scaled to long chain-of-thought training.

DAPO reports several pathologies:

  • entropy can collapse and exploration can shrink,
  • some prompt groups produce all-correct or all-wrong samples and therefore no useful relative gradient,
  • long responses can be weighted poorly,
  • truncated responses can receive noisy reward signals.

The DAPO paper introduces four core techniques.

Clip-Higher

The upper clipping range is relaxed relative to the lower range.

The goal is to make it easier for low-probability exploratory tokens to increase in probability, reducing premature entropy collapse.

Dynamic Sampling

If every sampled answer for a prompt is correct, or every answer is wrong, group-relative advantages can collapse to zero.

DAPO oversamples and filters those zero-gradient groups so that the effective training batch contains prompts with both successful and unsuccessful trajectories.

Token-Level Policy Gradient Loss

Instead of giving every completed sample equal weight after averaging within the sequence, DAPO changes the loss reduction so individual tokens in long reasoning traces contribute more appropriately.

This is especially relevant when long-CoT behavior grows during RL.

Overlong Reward Shaping

Hard truncation can punish an otherwise valid trajectory simply because it crossed a token limit.

DAPO uses length-aware shaping to reduce the reward noise caused by overlong and truncated samples.

11. The DAPO ablation is unusually informative

The paper reports a progressive AIME 2024 ablation on Qwen2.5-32B.

The sequence is:

  • Naive GRPO: 30
  • + Overlong Filtering: 36
  • + Clip-Higher: 38
  • + Soft Overlong Punishment: 41
  • + Token-level Loss: 42
  • + Dynamic Sampling: 50

The final DAPO system reaches 50 points on AIME 2024.

This number should not be compared directly with DeepSeekMath’s 51.7%, because the two results are on different benchmarks: DAPO’s number is AIME 2024, while DeepSeekMath’s 51.7% is MATH.

That distinction was blurred in the recorded source metadata and is separated here.

12. Verifiability is the hidden prerequisite of the whole pipeline

STaR, GRPO, and DAPO look different, but they share one hidden dependency.

The system needs a useful reward signal.

Mathematics works well because a final answer can often be checked automatically.

Code works well when unit tests, compilers, or execution environments provide machine-checkable feedback.

The situation is much harder for:

  • nuanced writing,
  • strategy,
  • diplomacy,
  • open-ended research judgment,
  • aesthetics,
  • values.

In those domains, a reward model may only be another fallible judge.

This leads to a strong general rule:

Train-time scaling is easiest when correctness is cheap to verify and hardest when “better” is itself ambiguous.

13. SFT and RL are not substitutes in a simple winner-take-all contest

The lecture contrasts supervised fine-tuning and reinforcement learning, but the practical relationship is complementary.

SFT is efficient when high-quality examples already exist.

RL becomes especially attractive when:

  • candidate solutions can be generated cheaply,
  • correctness can be verified,
  • and the space of possible trajectories is too large to label manually.

But RL still depends on the capabilities and data distribution established before the RL stage.

That is why “more RL” is not automatically better.

14. The strongest counterargument: the model may learn the reward, not the task

The entire self-improvement loop contains a familiar danger.

If the reward function is incomplete, the model can optimize the proxy instead of the intended objective.

If exploration collapses, the model may become confidently narrow.

If all sampled trajectories fail, group-relative learning may have no useful direction.

If the verifier accepts superficial patterns, the policy can learn those patterns.

DAPO itself is evidence for this concern: much of the paper is devoted not to inventing a new reasoning task, but to preventing the training dynamics from collapsing or becoming noisy.

So the strongest defensible conclusion is not:

“RL makes models recursively smarter without limit.”

It is:

RL can concentrate probability on better trajectories when the base model can generate them, the verifier can recognize them, and the optimization system preserves enough exploration to keep discovering useful variation.

15. What this article does not claim

  • The lecture did not show that RL can replace pretraining.
  • It did not show that STaR rationales are always causally faithful.
  • It did not show that GRPO is universally superior to PPO in every domain.
  • It did not show that DeepSeekMath’s 51.7% MATH result and DAPO’s 50-point AIME result are directly comparable.
  • It did not show that higher benchmark scores imply unlimited creation of new knowledge.
  • It did not show that verifiable-reward RL transfers cleanly to subjective domains.
  • It did not establish that longer reasoning traces are automatically better reasoning.

16. The real thesis of Part 6

Part 2 showed how inference-time compute can search over more possibilities.

Parts 3 and 4 showed that the quality of verification and feedback determines whether that search is useful.

Part 6 closes the loop.

Once a system can generate candidate trajectories and score them, the successful trajectories can be used to change future behavior.

STaR does it through iterative self-generated rationales.

DeepSeekMath does it through group-relative reinforcement learning.

DAPO shows how fragile that RL loop becomes when scaled to long reasoning — and how much engineering is required to keep exploration, reward signal, and sequence length under control.

The deeper transition is:

A self-improving system does not merely search for a better answer. It turns the evidence from that search into the distribution that will generate tomorrow’s answers.

Part 7 moves from this training loop into Self-Improvement and Deep Research Agents: what happens when the agent itself becomes the object of iterative improvement.

Claim Map

Speaker / course claims

  • Train-time scaling feeds successful reasoning back into model learning.
  • STaR, GRPO, and DAPO represent progressively more scalable self-improvement loops.
  • Verifiable tasks such as math and code are especially suitable for RL-based self-improvement.

Verified facts

  • STaR is arXiv:2203.14465 and iteratively generates, filters, rationalizes, and fine-tunes on model-generated reasoning.
  • DeepSeekMath is arXiv:2402.03300, continues pretraining DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens, introduces GRPO, and reports 51.7% on MATH plus 60.9% with 64-sample self-consistency.
  • DeepSeekMath reports stronger Maj@K improvement than Pass@K after RL in its tested setup.
  • DAPO is arXiv:2503.14476, introduces Clip-Higher, Dynamic Sampling, Token-Level Policy Gradient Loss, and Overlong Reward Shaping, and reports 50 points on AIME 2024 with Qwen2.5-32B.

Editorial interpretation

  • Train-time scaling is best understood as a system for converting verified search outcomes into persistent changes in future output probability.
  • The practical ceiling of self-improvement is jointly constrained by base-model reachability, verifier quality, and exploration stability.

Lecture Map

  • 00:00–00:16 — Train-time scaling, inference/training compute, verifiability
  • 00:16–00:41 — STaR and rationale bootstrapping
  • 00:41–00:53 — DeepSeekMath and GRPO
  • 00:53–01:06 — DAPO and large-scale RL stabilization
  • 01:06–01:12 — SFT vs RL and open questions

Primary Sources