Series: Self-Improving AI Agents — Part 5 of 9
Previous: Part 4 — Where Should an AI Agent Get Its Feedback?
Lecture: Stanford CS329A | Agent Evaluation
The previous lectures focused on how agents can search, verify, use tools, and learn from feedback.
Part 5 asks a different question:
How do we know whether an agent is actually good enough to trust with real work?
That question sounds like a benchmarking problem. The lecture argues that it is broader than that.
A model may score highly on short academic tests and still fail when a task lasts for an hour, requires context that was not written down, produces a professional deliverable rather than a text answer, or depends on retrieving and citing the right sources.
The lecture therefore moves through three evaluation regimes:
- METR: how long a task can the agent complete at a given reliability?
- GDPval: how good is the deliverable compared with work produced by industry professionals?
- DeepScholarBench: can the agent retrieve, synthesize, and cite live research accurately enough to support a real scholarly argument?
The larger thesis is:
capability, reliability, and economic usefulness are different variables.
1. Traditional benchmarks are increasingly too short
Many classic LLM evaluations ask a model to solve one bounded problem: answer a question, complete a proof, produce a short program, or select the correct option.
Those evaluations are still useful, but they compress the difficulty of real work into a short interaction.
Real agentic work looks different.
A task may require the system to plan, use multiple tools, inspect files, recover from mistakes, preserve state over many steps, notice when its first strategy has failed, and finally produce an artifact that another person can actually use.
The longer the trajectory, the more opportunities there are for small failures to compound.
This is why agent evaluation needs a notion of duration and reliability, not just correctness on isolated prompts.
2. METR asks: how difficult a task can an agent finish autonomously?
METR’s time-horizon work proposes a human-interpretable metric.
Instead of saying that a model gets 72% on benchmark X, ask:
How long would the corresponding task take a human expert, and at what task duration does the agent still succeed 50% of the time?
That duration is the agent’s 50% task-completion time horizon.
This definition is easy to misread. It does not mean the AI itself runs for that many minutes. It means the model is predicted to succeed half the time on tasks that take a human expert approximately that long to complete.
METR’s original time-horizon work found that this 50% horizon had been increasing exponentially across several years of frontier-model progress, with a historical doubling time of roughly seven months.
3. The lecture’s 59-minute number is a historical snapshot, not a current frontier number
The lecture uses Claude 3.7 Sonnet as a concrete example and cites a roughly one-hour 50% time horizon.
That was consistent with METR’s 2025 Claude 3.7 evaluation, which estimated the model’s 50% horizon at roughly 55–59 minutes depending on the evaluation setup.
But this number should not be presented as the current frontier in 2026.
METR continues to update its public time-horizon measurements as new models and task suites are evaluated. It also revised parts of its statistical methodology in 2026, including a correction that affected fitted time horizons—especially 80% horizons.
The safer conclusion is therefore about the shape of the result, not one frozen number:
frontier agents can handle progressively longer software tasks, but their reliable horizon remains substantially shorter than the tasks they can sometimes complete.
4. Reliability changes the answer dramatically
A 50% success rate is useful for measuring capability growth, but it is often unacceptable for deployment.
If an agent completes a one-hour task correctly only half the time, a human still needs to inspect or redo a large fraction of the work.
The lecture highlights this gap using the 2025 METR snapshot: when the required success probability rises from 50% to 80%, the effective task horizon falls sharply.
The exact numbers depend on model and methodology, and METR’s 2026 methodological updates make the historical figures unsuitable as immutable constants.
But the underlying engineering problem remains:
the frontier of “sometimes works” is much farther out than the frontier of “works reliably enough to delegate.”
For autonomous agents, that difference is often more important than the headline capability score.
5. Long tasks fail through error accumulation, not one dramatic mistake
The lecture groups common long-horizon failures into familiar categories:
- bad planning,
- wrong tool choice,
- reasoning or calculation errors,
- prematurely deciding the task is complete,
- getting trapped in repetitive loops.
Any one of these failures can be recoverable. The problem is that long tasks contain many decision points.
If each step is slightly unreliable, the probability of a fully correct trajectory can deteriorate quickly as the task becomes longer.
This is one reason short benchmark performance can coexist with poor long-horizon autonomy.
6. The “low-context contractor” analogy explains another failure mode
The lecture uses a useful analogy: an agent entering an unfamiliar codebase often behaves like an external contractor with very little organizational context.
The maintainer knows architecture, conventions, historical decisions, naming patterns, hidden dependencies, and which files are dangerous to touch.
The agent sees only what has been made explicit in its current context.
This creates a gap between technical capability and situational understanding.
The analogy should be treated as lecture framing rather than a universal measured constant. The deeper point is that tacit knowledge can dominate performance on real work.
An agent that has excellent coding skill but no model of the organization can still spend most of its effort rediscovering context that an experienced human already carries implicitly.
7. GDPval asks a harder question: is the deliverable economically useful?
METR measures task-completion difficulty through human completion time.
OpenAI’s GDPval moves to another axis: compare model-produced deliverables directly against work produced by experienced professionals.
GDPval v1 contains 1,320 specialized tasks across 44 occupations selected from nine major U.S. industries. The public gold subset contains 220 tasks. The tasks were created and reviewed by professionals with more than 14 years of experience on average.
These are not multiple-choice questions. Deliverables include documents, spreadsheets, slides, diagrams, multimedia, engineering materials, and other work products.
That distinction matters because professional quality includes more than factual correctness.
Formatting, aesthetics, completeness, file handling, instruction following, domain conventions, and integration of reference materials can all affect whether a result is usable.
8. GDPval shows roughly linear progress rather than the METR-style exponential horizon curve
OpenAI reports that frontier-model performance on GDPval improved roughly linearly over time.
In the released gold subset, Claude Opus 4.1 was the strongest evaluated model in the initial release and produced outputs judged as good as or better than the human comparison deliverable on 47.6% of tasks.
The same study found different model strengths: Claude Opus 4.1 was particularly strong on aesthetic and file-format-heavy deliverables, while GPT-5 was stronger on accuracy-oriented dimensions such as careful instruction following and calculations.
This contrast with METR is not a contradiction.
The two benchmarks are measuring different things.
- METR: how difficult a task can the agent complete?
- GDPval: does the final work product compete with an experienced professional’s deliverable?
A longer autonomous trajectory does not automatically produce professional-quality work.
9. GDPval also exposes the importance of context
The official GDPval description contains an important limitation: the first version is a one-shot evaluation.
It does not fully model the interactive process by which a professional asks clarifying questions, accumulates organizational context, negotiates priorities, receives feedback, and improves through multiple drafts.
That limitation is highly relevant to the lecture’s discussion of underspecified prompts.
Humans often infer missing context from organizational experience. Agents are much more dependent on what has been explicitly provided.
So better models help—but richer task context and better scaffolding also matter. OpenAI’s GDPval experiments report measurable improvements when models receive richer context and more reasoning effort.
10. Agent evaluation therefore needs to measure context acquisition, not just task execution
A realistic agent should not require every relevant fact to be perfectly encoded in the first prompt.
It should be able to notice uncertainty, ask for missing information, inspect artifacts, recover context from tools, and revise its own understanding.
This suggests an evaluation principle beyond the three benchmarks:
an agent should be evaluated on how well it acquires missing context, not only on how well it follows context that was already handed to it.
This is especially important for software, research, administration, and other domains where much of the work depends on tacit institutional knowledge.
11. DeepScholarBench asks whether deep research is actually verifiable
The third evaluation regime is DeepScholar-Bench, a live benchmark for generative research synthesis.
The task is deliberately realistic: given a recent research paper, generate its related-work section by searching the live literature, selecting important prior work, synthesizing the findings, and citing the sources.
The benchmark evaluates three broad dimensions:
- Knowledge synthesis: does the report organize and cover the important ideas?
- Retrieval quality: did the system find relevant and important papers?
- Verifiability: do the cited papers actually support the claims made in the report?
This is a much harsher test than generating fluent prose.
A system can write a convincing literature review while retrieving the wrong papers or attaching correct-looking citations to unsupported claims.
12. DeepScholarBench is designed to resist benchmark staleness
Static benchmarks have a contamination problem: once benchmark questions appear on the public web, future models may encounter them during training.
DeepScholar-Bench addresses this by using recent, high-quality arXiv papers and a continually updating data pipeline. The benchmark is designed to be refreshed as new literature appears.
That does not make contamination impossible, but it changes the evaluation from a frozen exam toward a moving research task.
This is important for deep-research systems because retrieval quality depends on the live state of the literature.
13. The “19% ceiling” in the lecture needs a version note
The lecture presents a striking result: existing deep-research systems remained below a 19% aggregate score in the benchmark snapshot it discusses.
The source history matters.
The early arXiv/Hugging Face version of DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis indeed states that no evaluated system exceeded 19% across all metrics.
However, the later OpenReview/ICLR revision reports the aggregate result differently: no system surpassed a geometric mean of 31% across the benchmark metrics.
This is not necessarily a contradiction; benchmark implementations, model sets, metric definitions, and aggregation can change between revisions.
The durable conclusion is:
deep-research systems remain far from saturating retrieval, synthesis, and citation-verification quality simultaneously.
14. The DeepScholar source ID in the lecture archive is also wrong
The source metadata recorded for this lecture lists DeepScholarBench as arXiv 2502.04357.
The actual paper is arXiv:2508.20033.
The public article therefore uses the verified paper and explicitly separates the lecture’s benchmark snapshot from later paper revisions.
This matters because an agent-evaluation article should not reproduce a citation-verification error while arguing that citation verification is essential.
15. The three benchmarks reveal three different bottlenecks
Placed side by side, the evaluations expose different failure surfaces.
METR — temporal reliability
Can the system maintain a successful trajectory as task duration grows?
GDPval — professional deliverable quality
Can the system produce something an experienced professional would judge competitive with real work?
DeepScholarBench — epistemic reliability
Can the system find the right evidence, synthesize it, and ensure that its claims are actually supported?
None of these dimensions can substitute for the others.
An agent can work for a long time and still produce a poor artifact. It can produce a polished artifact while misunderstanding the source material. It can retrieve good evidence but fail to finish a multi-hour workflow.
16. The strongest counterargument: benchmarks still simplify the workplace
Even these newer evaluations are controlled environments.
They do not fully capture organizational politics, changing priorities, access restrictions, ambiguous authority, contradictory stakeholders, long-term accountability, or the cost of repairing subtle mistakes after deployment.
GDPval explicitly notes that its first version is one-shot. METR’s current time-horizon suite is concentrated on software tasks. DeepScholar-Bench evaluates a specific kind of research synthesis.
So the correct conclusion is not that these benchmarks “measure real-world autonomy” in a complete sense.
They measure parts of it better than short-form academic benchmarks do.
the closer an evaluation gets to real work, the more evaluation itself becomes a systems-design problem.
17. What this article does not claim
- The lecture did not claim that a 50% time horizon means an agent can safely be left alone for that amount of real clock time.
- It did not claim that METR’s historical seven-month doubling trend is guaranteed to continue indefinitely.
- It did not claim that a near-50% GDPval score means half of professional jobs can already be automated.
- It did not claim that professional work can be reduced to one-shot deliverables.
- It did not claim that fluent research writing implies reliable source selection or citation accuracy.
- It did not establish that one benchmark can stand in for agent reliability as a whole.
18. The real thesis of Part 5
The progression of this course now becomes clearer.
Parts 2–4 asked how agents can search, verify, use tools, and learn from feedback.
Part 5 asks whether those improvements survive contact with longer tasks and real deliverables.
The answer is mixed.
Agents are clearly becoming capable of handling longer and more complex workflows. But three gaps remain:
- Reliability gap: a task the agent can sometimes complete is longer than a task it can reliably complete.
- Context gap: explicit prompts contain less knowledge than experienced workers carry implicitly.
- Verification gap: polished output can still be weak in evidence selection, citation support, or professional usability.
That leads to a stricter definition of progress:
an agent is not deployment-ready because it can solve the task. It is deployment-ready when it can solve the right task, with the right context, at the required reliability, and produce an output whose quality can be verified.
Part 6 will return from evaluation to learning and ask how capability itself can be expanded during training: what happens when we scale compute at train time using self-generated experience and reinforcement learning?
Claim map
- Speaker / course claims: agent evaluation must move beyond short-turn benchmarks toward time horizon, economic deliverable quality, and verifiable research synthesis; long tasks expose compounding error and context deficits.
- Verified facts: METR defines 50% and 80% task-completion time horizons using human-expert task duration and historically reported roughly seven-month doubling of the 50% horizon; GDPval contains 1,320 tasks across 44 occupations in nine U.S. industries, with a 220-task gold set and roughly linear model improvement over time; Claude Opus 4.1 reached 47.6% wins-plus-ties against expert deliverables in the initial gold-set evaluation; DeepScholar-Bench evaluates knowledge synthesis, retrieval quality, and verifiability using a live arXiv-based pipeline and is arXiv:2508.20033.
- Editorial interpretation: “capability, reliability, and economic usefulness are different variables” is the organizing synthesis of this article.
Lecture map
- 00:00–00:28 — agent evaluation, METR time horizons, reliability, and low-context work
- 00:28–00:52 — GDPval and expert-quality economic tasks
- 00:52–01:08 — DeepScholarBench and verifiable research synthesis
- 01:08–01:15 — long-tail reliability and future autonomous agents
Primary sources
- Stanford CS329A — Part 5 | Agent Evaluation
- Stanford CS329A — Self-Improving AI Agents
- METR — Task-Completion Time Horizons of Frontier AI Models
- METR — Measuring AI Ability to Complete Long Software Tasks
- OpenAI — GDPval
- DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis