Series: Self-Improving AI Agents — Part 4 of 9
Previous: Part 3 — When the Verifier Can Be Wrong
Lecture: Stanford CS329A | Learning from Feedback with Tools/Code
Part 3 ended with a difficult conclusion: model-based verification can improve search and learning, but the verifier itself can become a new source of error.
Part 4 changes the source of feedback.
Instead of asking another model to judge whether an answer looks correct, the agent can interact with something outside its own internal reasoning loop: a search environment, a code executor, a unit test, or a written constitution that constrains how AI feedback should be produced.
The lecture’s three core papers—ReAct, RLEF, and Constitutional AI—use very different mechanisms. But they all point toward the same systems principle:
self-improvement becomes more reliable when feedback is tied to an external constraint rather than generated only from the model’s own beliefs.
1. ReAct: reasoning becomes more useful when the world can answer back
ReAct—Synergizing Reasoning and Acting in Language Models—starts from a limitation of pure chain-of-thought reasoning.
A model can reason fluently using only its internal representations, but that reasoning is not automatically grounded in current facts or environment state. If the model begins from a wrong assumption, a long chain of thought can simply propagate the error.
ReAct interleaves:
- Thought: what should I do next?
- Action: call a tool or interact with the environment.
- Observation: receive new external information.
The next reasoning step is conditioned on the observation.
This changes the role of reasoning. It is no longer a closed internal monologue. It becomes a control layer that decides what information to request and how to respond to what the environment returns.
2. Why ReAct is different from both CoT and act-only agents
The ReAct paper explicitly compares reasoning-only and acting-only approaches.
Pure CoT can hallucinate because it relies on internal knowledge even when external evidence would be useful. Act-only approaches can interact with tools but lack an explicit reasoning trace for planning, exception handling, and updating strategy.
ReAct tries to combine both.
The original paper reports gains across HotpotQA, FEVER, ALFWorld, and WebShop. On ALFWorld and WebShop, ReAct outperformed imitation- and reinforcement-learning baselines by 34 and 10 percentage points in success rate, respectively, using only one or two in-context examples.
For HotpotQA and FEVER, the paper found that external Wikipedia interaction helped reduce hallucination and error propagation. It also found that hybrid strategies combining ReAct with chain-of-thought self-consistency could outperform either approach alone.
3. The deeper lesson: tools create a feedback loop, not just more context
It is tempting to describe tool use as “retrieval.” That is too narrow.
In ReAct, a tool call changes what the agent knows and therefore changes its next action. The sequence is adaptive:
reason → act → observe → update → act again
The important property is not simply that external information enters the prompt. It is that the agent’s policy can change after seeing the result.
This is the beginning of an environment-grounded learning loop.
4. ReAct still inherits the quality of its environment
External feedback is not automatically correct.
Search results can be incomplete. APIs can return stale values. Tools can fail. A website can be misleading. A robot sensor can be noisy.
So “grounded” does not mean “true.”
The more precise claim is:
external interaction can break the closed loop of model self-belief, but only if the environment provides useful signal.
This is why tool-using agents need retry logic, confidence estimation, fallback strategies, and sometimes cross-checking across tools.
5. Code offers something stronger than search: executable feedback
The second major paper in the lecture is RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.
Code is a particularly attractive domain because execution produces objective signals:
- does it compile?
- does it time out?
- does it crash?
- does it pass the tests?
- does a revised program improve on the previous one?
These signals are much harder for a model to fake than a natural-language explanation of why its code is correct.
RLEF trains models to use execution feedback over multiple steps rather than merely sampling independent programs.
6. The crucial RLEF idea is iterative repair
The RLEF paper focuses on competitive-programming tasks and asks whether a code model can learn to revise its own output after observing execution feedback.
Instead of treating each failed program as a dead end, the model receives the execution result and tries again.
This creates a loop:
generate code → execute → observe failure → revise → execute again
The paper reports state-of-the-art results with both 8B and 70B models and an order-of-magnitude reduction in the number of samples required compared with independent sampling baselines.
The important point is not the benchmark record by itself. It is that the model learns to use the error signal instead of treating every attempt as unrelated.
7. Why execution feedback is stronger than a language-model critique
A model can produce a plausible critique of broken code and still miss the bug.
A compiler error, failing unit test, or runtime timeout is different. It is generated by the execution environment, not by the same language model that wrote the code.
This creates a stronger form of grounding.
The feedback does not need to explain the entire solution. It only needs to constrain the next action.
Even a binary pass/fail signal can become useful if the agent is allowed to iterate and learn a repair policy.
8. But tests can be gamed too
Execution feedback feels objective, but it is only as good as the evaluator.
If the visible test suite is incomplete, a model can overfit to the tests. If a benchmark leaks information, the agent can exploit the harness instead of solving the underlying task.
This is why the lecture emphasizes separating feedback used for iterative repair from held-out evaluation signals.
The general principle is familiar from machine learning:
the feedback used to improve behavior should not be identical to the signal used to certify generalization.
The exact public/private-test mechanics should be read as one implementation pattern, not a universal definition of RLEF.
9. The RLEF source correction matters
The source metadata recorded for this lecture identifies the RLEF paper with arXiv 2402.04616. The actual paper is arXiv:2410.02089, titled RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.
The public article therefore uses the verified identifier and paper scope.
This is not a cosmetic correction. A source mismatch can silently attach claims from another paper to the wrong methodology.
10. Constitutional AI introduces a third kind of feedback: rule-conditioned judgment
The third research line is Anthropic’s Constitutional AI: Harmlessness from AI Feedback.
This system does not rely on a compiler or search engine. Instead, humans provide a set of written principles—a constitution—and AI models use those principles to critique and revise model outputs.
The paper has two major phases.
Supervised phase
The model generates an initial response, critiques that response according to a constitutional principle, revises it, and is then fine-tuned on the improved responses.
Reinforcement-learning phase
A model compares candidate responses according to the constitution. Those AI preferences are used to train a preference model, and the policy is optimized using reinforcement learning from AI feedback (RLAIF).
This shifts human labor from labeling thousands of outputs toward specifying higher-level principles.
11. Constitutional AI is externalized feedback—but not objective feedback
It is important not to collapse Constitutional AI into the same category as code execution.
A compiler can often produce a relatively objective failure signal. A constitution contains normative rules that still require interpretation.
The AI judge can misread the principle, apply it inconsistently, or inherit biases from the model used to interpret it.
So the three Part 4 feedback sources form a spectrum:
- Environment feedback: new facts or state from outside the model.
- Execution feedback: machine-checkable behavior from code and tests.
- Constitutional feedback: model judgments constrained by explicit human-written principles.
All three are more structured than unconstrained self-critique, but they are not equally objective.
12. The common architecture is a closed loop around the model
ReAct, RLEF, and Constitutional AI appear to solve different problems: factual grounding, code synthesis, and harmlessness.
Architecturally, however, they share the same pattern:
- the model produces an action or answer,
- an external process evaluates or transforms it,
- the result returns as feedback,
- the model changes its next output,
- the loop repeats or is distilled into training.
This is why Part 4 belongs in a course on self-improving agents.
The improvement does not come from “reflection” in the abstract. It comes from creating a channel through which the system receives information it did not already possess in the same form.
13. The strongest counterargument: external feedback can still become another reward to hack
A skeptical reading is necessary.
Agents do not care about our intention; they optimize the signal we give them.
If the search API is noisy, the agent may learn to over-trust bad sources. If the unit tests are incomplete, it may learn to target the tests. If the constitution is ambiguous, it may learn superficial patterns that satisfy the AI judge without satisfying the underlying norm.
So moving feedback outside the model is not enough.
The feedback channel itself must be designed to resist exploitation.
The strongest defensible conclusion is:
external feedback improves self-improvement when it is more informative, more falsifiable, and harder to game than the model’s internal self-evaluation.
14. What this article does not claim
- ReAct was not presented as universally superior to pure chain-of-thought on every task.
- Tool use was not presented as automatically truthful; environment feedback can be noisy or misleading.
- RLEF does not imply that code agents eliminate software engineers or solve arbitrary software repositories.
- Execution feedback is not immune to reward hacking when tests are incomplete.
- Constitutional AI does not remove all human input; humans still define the principles.
- RLAIF does not turn normative judgments into objective ground truth.
- None of the three methods establishes that one feedback loop is sufficient for every agentic task.
15. The real thesis of Part 4
Parts 2 and 3 focused on how to search and verify better. Part 4 asks a more foundational question:
Where does the corrective signal come from?
Three answers emerge.
- ReAct: let the environment answer back.
- RLEF: execute the artifact and use its behavior as feedback.
- Constitutional AI: turn human-written principles into scalable AI-generated feedback.
These approaches differ in objectivity, cost, and failure modes. But they all move the agent away from a closed loop in which one model generates, critiques, and trusts itself.
That is the deeper transition from language model to agent: the system becomes defined not only by the model, but by the feedback channels wrapped around it.
Part 5 shifts from feedback mechanisms to evaluation: once an agent can act and observe, how do we measure whether it is capable, reliable, and useful on longer tasks?
Claim map
- Speaker / course claims: tool interaction, execution feedback, and constitutional feedback are three major ways to create useful self-improvement loops around language models.
- Verified facts: ReAct interleaves reasoning traces and actions and reports strong results on HotpotQA, FEVER, ALFWorld, and WebShop; RLEF is arXiv:2410.02089 and reports state-of-the-art competitive-programming results with 8B and 70B models while reducing required samples by an order of magnitude; Constitutional AI uses self-critique/revision in a supervised phase and RLAIF in a reinforcement-learning phase with human oversight expressed through written principles.
- Editorial interpretation: “feedback quality is determined by how external, falsifiable, and hard to game the signal is” is the organizing synthesis of this article.
Lecture map
- 00:00–00:28 — ReAct, tool use, grounding, and environment feedback
- 00:28–00:46 — execution feedback and iterative code repair
- 00:46–01:01 — Constitutional AI, critique/revision, and RLAIF
- 01:01–01:11 — feedback loops, agent architecture, and open questions