Series: Self-Improving AI Agents — Part 1 of 9
Lecture: Stanford CS329A | Course Overview
For most of the modern large-language-model era, the dominant story of progress was straightforward: train on more data, spend more compute, increase model scale, and capabilities improve. Stanford’s CS329A course on Self-Improving AI Agents begins by accepting that history—and then asking what comes after it.
The first lecture is not a declaration that model scaling is over. It is a map of how the research agenda has widened. The instructors move from pretraining scaling laws to chain-of-thought prompting, post-training and RLHF, inference-time compute, reasoning models, agentic workflows, and finally the generator-verifier gap. The sequence matters because it shows how “self-improving agents” grew out of several earlier shifts rather than appearing as a separate technology category overnight.
The durable question underneath the lecture is this:
If making the model larger is only one way to improve performance, where else can we spend computation and supervision?
That question leads directly to test-time scaling, verification, tools, and learning loops—the subjects developed in the later lectures of the series.
1. The starting point: scaling really did work
The lecture begins with the success of pretraining scaling. This is important because the later turn toward inference-time compute makes little sense if we forget what came before it.
Large language models improved as researchers increased three tightly connected resources: model capacity, training data, and training compute. Aakanksha Chowdhery is particularly well placed to tell that story. She is the first author of the 2022 PaLM paper, which described a densely activated 540-billion-parameter Transformer trained across 6,144 TPU v4 chips. The paper reported strong few-shot performance across hundreds of tasks and large gains on multi-step reasoning benchmarks as scale increased.
That period established a powerful engineering intuition: if the loss curve continues to improve predictably with more scale, then larger training runs are not merely speculative bets. They are an industrial strategy.
But a scaling law does not imply that every additional dollar of pretraining produces the same economic return. The lecture frames the frontier problem as one of rising cost, finite high-quality data, and diminishing practical efficiency. The instructors do not say pretraining has stopped working. Their claim is narrower: the field increasingly needs additional axes of improvement.
2. Scale did more than reduce loss—it changed what prompting could unlock
The next step in the lecture is chain-of-thought prompting.
The 2022 paper Chain-of-Thought Prompting Elicits Reasoning in Large Language Models showed that providing intermediate reasoning demonstrations could substantially improve performance on arithmetic, commonsense, and symbolic reasoning tasks. One of its headline results used PaLM 540B: eight chain-of-thought exemplars were enough to produce state-of-the-art accuracy on GSM8K at the time.
For the course, this is an important transition. The model is no longer treated only as a fixed predictor whose quality is determined at training time. The way inference is organized begins to matter.
However, the language of “emergent abilities” deserves caution. The lecture presents chain-of-thought as an ability that becomes useful only once models are sufficiently capable. That framing is consistent with influential work on emergence. But the scientific interpretation is contested. A 2023 paper, Are Emergent Abilities of Large Language Models a Mirage?, argued that some apparently discontinuous capability jumps can arise from the choice of evaluation metric rather than a literal phase transition in model behavior.
So the verified fact is that larger models in the cited experiments benefited much more from chain-of-thought prompting. The stronger claim—that intelligence suddenly “appears” at a particular parameter threshold—is better treated as an open research interpretation.
3. Post-training changed the meaning of “better model”
The lecture then shifts from scale to post-training.
Pretraining teaches a model to continue text. A deployed assistant has a different job: understand a request, follow instructions, avoid obvious harms, produce useful answers, and behave consistently with human preferences. Those objectives are not automatically optimized by next-token prediction.
OpenAI’s InstructGPT work made this distinction concrete. The training pipeline combined supervised fine-tuning on human-written demonstrations with rankings of model outputs, a learned reward model, and reinforcement learning from human feedback. In the paper’s human evaluations, a 1.3B-parameter InstructGPT model was preferred over the 175B GPT-3 model on the prompt distribution used in the study.
That result does not mean “small models beat large models” in general. It means that training objective and feedback can matter enough to overturn a large parameter-count advantage on a specific deployed behavior metric.
This is the conceptual bridge to self-improvement. If feedback can change what a pretrained model does, then the next question is where that feedback comes from, how cheaply it can be generated, and whether the model can eventually participate in producing it itself.
4. The new axis: spend more compute after training
Around the 20-minute mark, the lecture introduces the idea that becomes the foundation for Part 2 of the series: inference-time or test-time scaling.
Instead of asking one model for one answer, generate many attempts. If at least some of those attempts are correct, a verifier may be able to identify a good solution.
The 2024 paper Large Language Monkeys: Scaling Inference Compute with Repeated Sampling, coauthored by Azalia Mirhoseini, tested this idea across coding, formal proofs, and math tasks. The researchers generated up to 10,000 samples per problem in several settings and found that solution coverage—whether at least one correct solution appeared—continued increasing with sample budget.
The most important result is not “10,000 samples always beats a frontier model.” The paper is more nuanced.
- On coding and formal-proof tasks, where solutions can be checked automatically, repeated sampling can turn higher coverage directly into higher task performance.
- On MATH with Llama-3-8B-Instruct, coverage rose from 79.8% at 100 samples to 95.3% at 10,000 samples.
- On SWE-bench Lite, DeepSeek-Coder-V2-Instruct increased from 15.9% solved with one sample to 56% with 250 samples.
- But in domains without reliable automatic verification, majority voting and reward-model selection plateaued as sample counts grew.
This is the first major systems lesson of the course: generation can scale faster than selection.
A model may be capable of producing a correct answer somewhere in a large sample set, yet the system may still fail if it cannot recognize that answer.
5. Why a verifier changes the economics of intelligence
The lecture repeatedly returns to what it calls the generator-verifier gap.
Some domains are unusually favorable for self-improvement because correctness can be checked cheaply. Code can be compiled and tested. Formal proofs can be passed through a proof checker. Mathematical answers sometimes have exact or structured validation rules.
Other domains are harder. A persuasive strategy memo, a scientific hypothesis, a policy analysis, or a long-form article may not have a cheap binary test. In those cases, the evaluator may be another model or a human expert—and both can be expensive or fallible.
This distinction explains why coding and mathematics appear so often in current agent research. It is not simply because researchers prefer those domains. They offer unusually dense feedback.
Self-improvement becomes easier when the environment can answer the agent with something stronger than “this looks good.”
6. Reasoning models combine search with learning
The lecture next discusses reasoning models such as OpenAI’s o-series, DeepSeek-R1, and “thinking” systems. Here the evidence boundary becomes important.
The instructors describe a general flywheel: spend more compute during inference, generate longer reasoning trajectories, identify successful trajectories, and feed useful data back into post-training. This high-level pattern is supported by public research across reinforcement learning and synthetic-data systems.
But the exact internal training pipelines of proprietary frontier models are not fully public. It would therefore be too strong to present every specific mechanism in the lecture as independently verified fact about o1, o3, or other closed systems.
The safer conclusion is architectural: modern reasoning research increasingly treats inference and training as connected stages rather than isolated phases. Search during inference can produce data. Feedback can label that data. Training can make the next round of inference better.
That is the beginnings of a self-improvement loop.
7. From chatbot to agent: the model now has to do things
At roughly 41 minutes, the lecture changes level again. The topic is no longer only how a model answers a question. It is how a system completes a goal.
The instructors distinguish a chatbot from an agent by adding several ingredients:
- a goal that persists across multiple steps,
- access to an environment such as files, a terminal, the web, or APIs,
- tools that can change that environment,
- memory or state that survives across steps,
- feedback from tool execution,
- the ability to revise a plan after failure, and
- a stopping rule that determines when the goal is complete.
This is a more demanding standard than fluent conversation. An agent must convert language into action and then incorporate what actually happened.
The lecture walks through orchestration patterns that have become common in agent systems: prompt chaining, routing, parallelization, orchestrator-worker structures, and evaluator or judge loops. These patterns are less glamorous than a new model release, but they matter because they determine how a model is embedded into a reliable process.
8. Coding agents expose the difference between generation and completion
Software engineering provides the clearest example.
A coding model can generate a patch in seconds. An agent must do more: inspect a repository, understand the task, edit files, run tests, interpret failures, revise the patch, and decide whether the task is actually complete.
The environment creates a sequence of correction signals. A failing test is not just a negative score—it contains information about what changed and what still needs fixing.
This makes the coding-agent loop a concrete instance of the broader architecture introduced in the lecture:
goal → action → observation → verification → revision → completion
The key constraint is still verification. If the tests are incomplete, the agent may optimize to the test suite while remaining wrong in production. If the task specification is vague, a technically correct patch may solve the wrong problem.
That leads to another point the instructors emphasize: intent clarification is part of agent capability. A system that executes an ambiguous instruction immediately may be less useful than one that recognizes missing requirements and asks the right question first.
9. The strongest counterargument: perhaps the field is still mostly riding model scaling
A fair reading of the lecture needs a serious objection.
Test-time scaling, agent orchestration, verifiers, and post-training all depend on the quality of the base model. A weak generator cannot search its way to solutions that are effectively absent from its distribution. Better models also tend to use tools more reliably, understand longer context, and generate stronger candidate plans.
There is also a measurement problem. What looks like a new “reasoning ability” can sometimes reflect prompting, metric choice, benchmark contamination, or better extraction of a capability that was already present.
So the first lecture should not be summarized as “bigger models are finished.” The more defensible interpretation is:
Pretraining scale remains foundational, but it is no longer the only place where additional compute, data, and feedback can produce capability gains.
That is a much more consequential claim because it changes system design. Engineers can now ask not only how large the model should be, but how much compute should be spent at inference, how candidates should be verified, which tools should be available, what feedback should be retained, and what part of that experience should flow back into training.
10. What this article does not claim
The lecture transcript is especially useful here because it preserves several boundaries that are easy to lose in a shorter summary.
- They did not claim that pretraining scaling is obsolete.
- They did not claim that reasoning models think in a human or conscious sense.
- They did not claim that current agents can replace all human software work autonomously.
- They explicitly acknowledged the cost of repeated sampling and the difficulty of verification in open-ended domains.
- They treated the contribution of pretraining versus post-training/RL to reasoning capability as an open research question rather than a settled result.
These qualifications are not footnotes. They define the difference between a useful research roadmap and a hype narrative.
11. The real thesis of Part 1
The first lecture starts with scaling laws, but its destination is systems engineering.
The progression is:
- Scale the pretrained model to obtain broad capabilities.
- Use prompting and post-training to make those capabilities easier to elicit and align with user intent.
- Spend compute during inference to search a larger solution space.
- Build verifiers that can distinguish useful trajectories from convincing errors.
- Connect the model to tools and environments so actions produce external feedback.
- Feed successful behavior back into learning where possible.
- Evaluate the whole loop rather than judging the model from one answer.
That sequence is the conceptual foundation for the rest of the series.
Part 2 will narrow the lens to one question that Part 1 only introduces: how far can performance scale if we keep the model fixed and spend more computation at test time?
Claim map
- Speaker / course claims: the research frontier is expanding from pretraining-only scaling toward inference-time compute, post-training, verification, and agentic feedback loops; robust verifiers are a major bottleneck; intent clarification matters in real workflows.
- Verified facts: PaLM was a 540B densely activated model; the cited CoT paper found large gains on reasoning benchmarks with sufficiently capable models; InstructGPT used human demonstrations, rankings, a reward model, and RLHF; repeated sampling increases solution coverage and is especially useful when outputs can be automatically verified.
- Editorial interpretation: Part 1 is best read as a transition from “model-centric scaling” to “system-level scaling.” That phrase is the interpretation of this article, not a direct quotation from the instructors.
Lecture map
- 00:02–00:06 — pretraining scaling laws and the history of model scale
- 00:06–00:11 — chain-of-thought and the emergence debate
- 00:12–00:19 — instruction tuning, human preference data, reward models and RLHF
- 00:19–00:28 — inference-time scaling and repeated sampling
- 00:28–00:41 — reasoning models and self-improvement loops
- 00:41–00:53 — chatbot-to-agent transition and orchestration patterns
- 00:49–00:53 — generator-verifier gap
- 00:53–00:58 — intent clarification and practical applications
Primary sources
- Stanford CS329A — Part 1 | Course Overview
- Stanford CS329A — Self-Improving AI Agents
- PaLM: Scaling Language Modeling with Pathways
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Training Language Models to Follow Instructions with Human Feedback
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Are Emergent Abilities of Large Language Models a Mirage?