Test-Time Compute Is a Resource Allocation Problem — Stanford CS329A Part 2

Series: Self-Improving AI Agents — Part 2 of 9
Previous: Part 1 — Why AI Progress Is Moving Beyond Bigger Models
Lecture: Stanford CS329A | Test-Time Compute Scaling

한국어판

Part 1 introduced a new axis of AI progress: instead of putting all additional compute into training a larger model, spend some of it after training, when the model is solving a specific problem.

Part 2 asks the harder question. How should that test-time compute actually be spent?

The naïve answer is “generate more samples.” Stanford’s second CS329A lecture starts there, but does not end there. It moves through repeated sampling, automated verification, process reward models, adaptive compute allocation, and finally Archon—a framework that treats the inference pipeline itself as something that can be designed and searched.

The central engineering lesson is more precise than “more inference is better”:

Test-time compute only becomes a reliable scaling axis when the system knows where to spend it and how to judge what it produced.

1. Repeated sampling reveals hidden solution coverage

The lecture opens by revisiting Large Language Monkeys, the 2024 paper coauthored by Azalia Mirhoseini.

The setup is deliberately simple. Keep the model weights fixed. Give the same problem multiple independent attempts. Measure coverage: the fraction of problems for which at least one generated sample is correct.

Across several models and tasks, coverage continued increasing as the sample budget grew across orders of magnitude. On SWE-bench Lite, DeepSeek-Coder-V2-Instruct rose from 15.9% of issues solved with one sample to 56% with 250 samples. In coding and formal proof domains, where generated outputs can be checked automatically, that additional coverage can be converted directly into better system performance.

This result changes what “model capability” means. A single answer measures one draw from a distribution. Repeated sampling asks a different question: does the model know enough to produce a correct trajectory at all?

Those are not the same metric.

2. But coverage is not the same as usable accuracy

The same paper exposes the immediate limitation.

Suppose a model generates 1,000 answers and only two are correct. Coverage says the model is capable of solving the problem. A deployed system still needs to identify the two good answers.

In domains without reliable automatic verifiers, common selection methods such as majority voting and learned reward models eventually plateau. More samples keep creating more possibilities, but the selector stops improving at the same rate.

This is the generator-verifier gap in operational form:

the model may be able to generate the answer before the system is able to recognize it.

That distinction explains why test-time scaling works especially well in code, formal proof, and other domains where the environment can reject wrong outputs cheaply.

3. Why the benchmark curve can look like a power law

The lecture discusses an inference scaling law: as the number of attempts increases, aggregate benchmark coverage often follows a smooth power-law-like curve.

There is an important source distinction here.

The 2024 Large Language Monkeys paper established the empirical scaling pattern across tasks and models. A 2025 follow-up paper, How Do Large Language Monkeys Get Their Power (Laws)?, gave a more explicit mathematical explanation for why aggregate power-law behavior can emerge even though success on an individual problem improves exponentially with repeated independent attempts.

For one problem with single-attempt success probability p, the probability of succeeding at least once after k independent samples is:

1 − (1 − p)k

That is exponential in k. The follow-up work argues that benchmark-level power-law behavior can arise when the distribution of per-problem success probabilities has a heavy tail: a small set of extremely difficult problems dominate the aggregate failure curve.

The practical implication is more important than the algebra. Benchmarks are not homogeneous. Easy problems saturate quickly. Hard problems absorb most of the additional compute.

4. Not every query deserves the same compute budget

That observation leads directly to Charlie Snell and colleagues’ 2024 paper, Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

The paper asks a resource-allocation question: given a fixed inference budget, how should compute be allocated to maximize performance?

The answer is not “use the same strategy everywhere.” The authors find that the effectiveness of different test-time techniques depends strongly on prompt difficulty. Their compute-optimal strategy improves efficiency by more than 4× over a best-of-N baseline. In FLOPs-matched experiments, on problems where the smaller base model already had a non-trivial chance of success, test-time compute could outperform a model 14× larger.

That result needs careful interpretation.

It does not mean a small model plus extra inference always beats a much larger model. The condition matters: the smaller model must already have some useful probability mass on good solutions. On the hardest problems, a stronger base model can still be the better investment.

5. Search strategy should adapt to difficulty

The lecture uses the Snell et al. work to contrast two broad ways of spending test-time compute.

  • Parallel exploration: sample many independent candidate solutions.
  • Adaptive improvement: use feedback to update or refine the response distribution as the model works on the problem.

For relatively easy problems, repeatedly revising a promising trajectory may be efficient. For harder problems, broader exploration can be more valuable because a single trajectory may get stuck in a bad region of the search space.

This suggests a more general architecture principle:

compute should be routed by uncertainty and difficulty, not allocated uniformly.

A production agent that spends the same number of tokens, tool calls, and verifier passes on every request is likely wasting compute on easy cases and under-investing in hard ones.

6. Process feedback can prune the search before the final answer

Another design choice is where verification happens.

An outcome-based evaluator looks at the completed answer and scores the result. A process-based evaluator scores intermediate reasoning steps.

That distinction matters for search. If a branch becomes implausible at step three, a process verifier can prune it before the system spends another thousand tokens completing the rest of the trajectory.

In tree- or beam-search-style inference, this creates a mechanism for spending compute selectively: expand promising branches, cut weak ones, and reserve additional budget for ambiguous cases.

This is also why Part 2 naturally leads to Part 3. Scaling search only helps if the verifier is robust enough to guide the search without being gamed.

7. Automated verification defines where test-time scaling is easiest

The lecture surveys domains with different verifier quality.

Formal proof systems can provide machine-checkable correctness. Software tasks can use tests. GPU-kernel generation can compare candidate outputs with a trusted reference implementation.

KernelBench is a useful example. The benchmark asks models to generate optimized GPU kernels corresponding to reference PyTorch programs. Correctness can be checked by comparing candidate outputs with the reference, while performance can be benchmarked separately.

But even here, “automatic” does not mean infallible. The KernelBench maintainers explicitly warn about reward hacking and suspiciously good benchmark results. If the test harness or correctness checks are incomplete, an optimizer can exploit the evaluator rather than solve the intended problem.

So the deeper rule is not “use automatic tests.” It is:

the ceiling of test-time scaling is set by the quality of the feedback channel.

8. Archon turns inference into an architecture-design problem

The final third of the lecture moves beyond repeated sampling.

What if inference is not one model call repeated N times, but a multi-layer system?

The Archon paper formalizes that idea. Instead of relying on a single LLM, Archon combines inference-time components such as generators, critics, rankers, fusers, verifiers, and model-based unit-test modules. The architecture can stack those components into layers and use inference-time architecture search to optimize the composition for a benchmark and compute budget.

The paper describes Archon as a way to turn system design into a hyperparameter-search problem. Which models should generate candidates? How many samples? Should responses be critiqued before ranking? Should several candidates be fused into a new answer? How many layers are worth the cost?

Those are inference questions, not training questions.

9. Fusion is different from selecting the best candidate

One of Archon’s more interesting findings is that the best final response does not always have to exist among the original candidates.

A ranker chooses among existing answers. A fuser receives several candidates and generates a new response that attempts to combine their strongest elements.

The Archon paper reports that candidate fusion can improve final response quality beyond the oracle best individual candidate in some experiments. That is conceptually important: extra inference compute is not limited to searching a fixed set. It can also be used to transform and synthesize candidate information.

In other words, inference can have depth as well as width.

10. The verified Archon result is strong—but narrower than the lecture summary

The primary paper is important here because the source metadata recorded for this lecture contains two mismatches.

The correct arXiv identifier for Archon: An Architecture Search Framework for Inference-Time Techniques is 2409.15254. Its abstract reports that the best Archon architectures, using all available LLMs, improved average accuracy by 15.1 percentage points over the compared frontier single-call models and prior inference architectures across the evaluated benchmarks.

When restricted to open-source LLMs, the paper reports an average advantage of 11.2 percentage points over single-call state-of-the-art LLMs in its evaluation.

That is still a substantial result, but it is different from saying “open-source models beat GPT-4o and Claude 3.5 Sonnet by 14.1% on average.” The public article should preserve the paper’s actual scope.

11. More inference has a different cost structure from more training

A student raises one of the strongest economic objections in the lecture.

Pretraining is expensive, but once the model exists, the cost is amortized across every future query. Test-time compute is a variable cost paid again for each problem.

That matters.

If every query requires hundreds of samples, multiple critics, rankers, and fusers, inference can become prohibitively expensive and slow. Real-time chat and high-volume consumer services face very different economics from offline theorem proving, code migration, scientific search, or high-value enterprise decisions.

This means the comparison between training scale and inference scale cannot be reduced to benchmark accuracy. It is an economic allocation problem involving:

  • fixed training cost,
  • per-query inference cost,
  • latency constraints,
  • task value,
  • verifier quality, and
  • how often the same capability will be reused.

Test-time scaling is especially attractive when retraining is impossible, when the model is available only through an API, or when additional compute can be reserved for a small number of high-value tasks.

12. The strongest counterargument: this may be an expensive way to recover what a better model already knows

A skeptical interpretation is straightforward.

If a model needs 1,000 attempts, a process verifier, several critics, and a fusion layer to match a larger model’s one-shot output, perhaps the cleaner solution is simply to use the stronger model.

Sometimes that is correct.

Test-time scaling cannot manufacture a solution that the base models almost never generate, and complex inference pipelines introduce latency, engineering complexity, and new failure modes. Verifiers can be wrong. Generated unit tests can be incomplete. Fusion can amplify shared errors instead of canceling them.

The strongest evidence in favor of test-time compute is therefore conditional rather than universal:

when a base model already has a non-trivial chance of producing useful components, and when feedback is reliable enough to guide search, extra inference compute can be a highly effective substitute for some additional model scale.

That is very different from claiming that inference scaling replaces pretraining.

13. What this article does not claim

  • The lecture did not claim that pretraining has become unnecessary.
  • It did not claim that majority voting is useless on every task; its failure is most severe when correct answers are rare.
  • It did not claim that more samples automatically produce better deployed accuracy without a selector or verifier.
  • It did not establish that one inference architecture is universally optimal across all domains and budgets.
  • It did not eliminate the economics of latency and per-query cost.
  • Archon itself uses domain-informed design constraints and search; architecture search is not a proof that human system design is obsolete.

14. The real thesis of Part 2

The lecture begins with “sample more,” but ends with a much richer view of inference.

Test-time compute can be spent on:

  1. generating more alternatives,
  2. revising promising trajectories,
  3. verifying intermediate steps,
  4. pruning weak branches,
  5. ranking candidate answers,
  6. fusing partial solutions, and
  7. searching over entire inference architectures.

The hard problem is not whether more compute can help. It can. The hard problem is where to spend the next unit of compute.

That makes inference scaling less like “letting the model think longer” and more like designing a compute scheduler for intelligence.

Part 3 will address the constraint that now becomes unavoidable: if verification controls the search, how do we make the verifier itself robust?

Claim map

  • Speaker / course claims: test-time compute is a major scaling axis; problem difficulty should influence how inference budget is allocated; the generator-verifier gap limits repeated sampling; multi-layer inference architectures can outperform single-call systems.
  • Verified facts: repeated sampling increases coverage; DeepSeek-Coder-V2-Instruct improved from 15.9% to 56% on SWE-bench Lite in the cited study; Snell et al. report >4× efficiency over best-of-N and a FLOPs-matched result against a 14× larger model under stated conditions; Archon combines generation, ranking, fusion, critique, verification, and unit testing and reports +15.1 percentage points using all available LLMs and +11.2 points with open-source LLMs in its evaluated setting.
  • Editorial interpretation: “test-time scaling is a compute-allocation problem” is the organizing thesis of this article. It synthesizes the lecture and cited papers rather than quoting one source verbatim.

Lecture map

  • 00:00–00:05 — inference as a third scaling stage; repeated sampling review
  • 00:05–00:12 — inference scaling curves and hard-problem tails
  • 00:12–00:19 — automated verification and the generator-verifier gap
  • 00:19–00:27 — discussion of verifier design and failure modes
  • 00:27–00:41 — compute-optimal test-time scaling and process-based verification
  • 00:41–00:46 — parallel versus tree-style search
  • 00:46–01:03 — Archon, fusion, deep inference layers, and architecture search

Primary sources