오래 일할 수 있는 AI와 믿고 맡길 수 있는 AI는 다르다 — Stanford CS329A Part 8

시리즈: 자가 개선 AI 에이전트 — 8/9

이전 편: Part 7 — Search and Deep Research Agents

강의: Stanford CS329A | Agentic Evaluations and Long-Horizon Tasks

원본 영상: https://www.youtube.com/watch?v=8JAqLnTaZu4

English version

1. Part 5 이후에도 Part 8이 필요한 이유

Part 5는 이미 중요한 구분을 다뤘다.

Capability와 reliability는 같은 것이 아니다.

Part 8은 이 문제를 실제 deployment 관점으로 더 좁힌다.

질문은 이제 단순히 “이 모델이 이 문제를 풀 수 있는가?”가 아니다.

다음 세 질문으로 바뀐다.

  • 얼마나 긴 task를 맡길 수 있는가?
  • 얼마나 높은 확률로 성공하는가?
  • 결과물이 실제로 돈을 지불할 만큼 가치 있고 검증 가능한가?

강의는 이를 세 benchmark family로 나눈다.

  1. METR: task-completion time horizon과 reliability
  2. GDPval: professional deliverable quality와 economic usefulness
  3. DeepScholar-Bench: research synthesis, retrieval, citation verifiability

긴 autonomy, 높은 reliability, 경제적 가치, research verifiability는 서로 다른 변수다.

2. METR은 평가 단위를 질문에서 과제로 바꾼다

전통적인 benchmark는 “질문에 맞는 답을 냈는가?”를 본다.

METR은 특정 성공 확률에서 어느 정도 길이의 task까지 agent가 완수할 수 있는가를 본다.

핵심 metric은 time horizon이다.

하지만 이 표현은 오해하기 쉽다.

60분 time horizon은 AI가 실제 wall-clock으로 60분 동안 autonomous하게 일했다는 뜻이 아니다.

METR의 task distribution에서 인간 전문가가 약 60분 걸리는 난이도의 task를 특정 확률로 성공할 것으로 예측한다는 뜻이다.

즉 duration은 human-equivalent task difficulty다.

3. 약 7개월 doubling은 실제 연구 결과지만 법칙은 아니다

METR의 초기 연구는 frontier agent의 50% task-completion time horizon이 과거 약 6년 동안 대략 7개월마다 두 배 증가했다고 보고했다.

Claude 3.7 Sonnet의 50% horizon은 evaluation setup에 따라 약 55~59분이었다.

이 결과가 중요한 이유는 짧은 benchmark가 보지 못하는 multi-step task capability의 증가를 보여주기 때문이다.

하지만 이를 물리 법칙처럼 외삽하면 안 된다.

METR은 이후 task suite와 tooling을 업데이트했고, 현재는 16시간을 넘는 측정은 current task suite에서 unreliable하다고 명시한다.

Agent time horizon은 강한 역사적 exponential trend를 보여왔지만, 정확한 slope와 미래 도달 시점은 methodology에 따라 달라질 수 있다.

4. 실제 deployment에서 더 중요한 것은 50%가 아니라 reliability tail이다

50% success probability는 capability trend를 측정하기에는 유용하다.

하지만 production 기준으로는 낮다.

1시간짜리 일을 절반 확률로 실패하는 agent는 여전히 많은 supervision과 checking이 필요하다.

Higher-reliability time horizon은 50% horizon보다 훨씬 짧다.

Agent가 시도할 수 있는 task의 길이는 빠르게 증가하지만, 인간이 감독 없이 맡길 수 있는 task의 길이는 그보다 짧다.

이 reliability gap이 실제 deployment의 핵심 병목이다.

5. High-reliability metric 자체도 methodology에 민감하다

METR의 2026 methodology analysis는 high-reliability time horizon이 fitted success-curve 가정에 상당히 민감할 수 있음을 보여준다.

Language model은 매우 짧고 쉬운 task에서도 가끔 실패한다.

그런데 statistical model이 “task가 짧아질수록 거의 반드시 성공한다”고 강하게 가정하면 80% 같은 reliability tail을 잘못 추정할 수 있다.

따라서 “80% horizon은 정확히 X분” 같은 숫자는 versioned estimate로 봐야 한다.

6. METR의 low-context limitation은 headline보다 중요하다

METR은 human task duration baseline이 low-context skilled contractor, new hire, freelancer에 더 가깝다고 명시한다.

실제 조직 업무는 다음과 같은 context에 크게 의존한다.

  • undocumented history
  • internal convention
  • previous conversation
  • tacit knowledge
  • 기존 codebase familiarity
  • 조직 내 권한과 책임 관계

따라서 METR의 2시간 task를 “high-context 직원의 실제 업무 2시간을 그대로 대체할 수 있다”고 읽으면 안 된다.

7. GDPval은 완료 여부보다 결과물의 가치를 평가한다

METR이 task duration을 본다면 GDPval은 다른 질문을 한다.

경험 많은 전문가가 이 deliverable을 인간 전문가의 결과물과 비교했을 때 실제로 comparable하다고 판단하는가?

GDPval v1은 공식적으로:

  • 미국 GDP 기여 상위 9 sectors
  • 44 occupations
  • full set 1,320 tasks
  • open-source gold subset 220 tasks
  • task 작성·검토 전문가 평균 경력 14년 이상

으로 구성된다.

평가 대상도 시험 문제가 아니라 실제 professional work product다.

Legal document, engineering plan, spreadsheet, support artifact, care plan, presentation, diagram 같은 산출물을 요구한다.

8. GDPval은 320-task benchmark가 아니다

강의 자료에서는 GDPval을 약 320개 task로 기술한다.

공식 초기 버전은 1,320 tasks다.

Gold subset은 220 tasks다.

또 강의 자료에 제시된 GDPval arXiv ID도 잘못돼 있다.

검증된 실제 논문은:

arXiv:2510.04374 — GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

이다.

9. GDPval 초기 약 50% frontier 결과는 역사적 snapshot이다

GDPval 초기 release에서는 frontier model이 professional deliverable quality에 상당히 접근했다.

그 release의 gold-set evaluation에서 Claude Opus 4.1은 human expert deliverable과 비교해 wins+ties 기준 절반에 가까운 수준을 기록했다.

하지만 이 수치를 2026년 현재 ceiling처럼 쓰면 안 된다.

이후 같은 GDPval family에서 OpenAI는:

  • GPT-5.2 Thinking 70.9% wins-or-ties
  • GPT-5.4 83.0% wins-or-ties

를 공개했다.

Professional deliverable quality는 계속 빠르게 개선되고 있지만, score를 비교할 때 model version, benchmark setup, reasoning effort, scaffolding을 함께 기록해야 한다.

10. Faster and cheaper는 autonomous와 같은 뜻이 아니다

Raw model inference는 전문가보다 매우 빠르고 저렴할 수 있다.

하지만 실제 workplace에서는 task specification, reference file preparation, review, retry, correction, integration, accountability 비용이 남는다.

GDPval 논문도 human oversight와 repair를 포함한 workflow를 별도로 분석한다.

실제 경제성 질문은:

AI generation 이후에도 human time이 얼마나 남는가?

다.

11. DeepScholar-Bench는 그럴듯한 연구와 검증 가능한 연구를 분리한다

세 번째 benchmark는 professional deliverable이 아니라 research synthesis를 본다.

DeepScholar-Bench는 최근 literature를 이용해 논문의 Related Work를 생성하게 한다.

핵심 평가축은 세 가지다.

  1. Knowledge synthesis
  2. Retrieval quality
  3. Verifiability

필수 foundational literature를 찾고, 핵심 technical contribution을 뽑고, paper 간 관계를 정확히 설명하며, citation이 실제 claim을 support해야 한다.

12. DeepScholar-Bench는 arXiv:2508.20033이다

강의 자료에 제시된 2501.09876은 잘못된 ID다.

검증된 실제 논문은:

arXiv:2508.20033 — DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

이다.

초기 arXiv version에서는 어떤 system도 모든 metric에서 19%를 넘지 못했다고 보고했다.

하지만 later ICLR revision은 geometric mean 31%를 넘긴 system이 없었다고 보고한다.

이 둘을 19→31의 단순 performance trend로 해석하면 안 된다.

Benchmark version, evaluated system, aggregation이 바뀌었기 때문이다.

Generative research synthesis는 아직 saturation과 거리가 멀다.

13. Fluency trap은 agent evaluation에서 특히 위험하다

Research agent output은 매우 그럴듯해 보일 수 있다.

문장은 매끄럽고, structure는 정돈돼 있고, citation도 많다.

하지만 이것만으로 factual grounding이 보장되지는 않는다.

DeepScholar-Bench의 가치는 presentation quality와 retrieval/verifiability를 분리한다는 데 있다.

Output이 더 유려해질수록 evidence 검증의 중요성이 커진다.

14. 세 benchmark는 서로 다른 failure surface를 측정한다

METR

질문: 특정 reliability에서 어느 정도 난이도의 task를 완수하는가?

핵심:

  • duration
  • autonomy
  • failure probability
  • low-context robustness

GDPval

질문: 인간 전문가와 비교할 때 deliverable의 품질과 경제적 가치는 어떠한가?

핵심:

  • professional quality
  • usefulness
  • multimodal artifact
  • economic relevance

DeepScholar-Bench

질문: evidence를 제대로 찾고, 종합하고, citation으로 검증할 수 있는가?

핵심:

  • retrieval
  • synthesis
  • grounding
  • verifiability

한 benchmark의 강한 score가 다른 두 benchmark의 강한 score를 보장하지 않는다.

15. 가장 강한 반론: benchmark가 deployment readiness보다 benchmark environment 적응력을 측정하는 것 아닌가?

세 benchmark 모두 현실을 단순화한다.

METR task는 주로 self-contained하고 well-specified된 digital task다.

GDPval v1은 largely one-shot evaluation이며 반복 협업, 변화하는 priority, stakeholder negotiation, 조직 politics를 완전히 담지 못한다.

DeepScholar-Bench는 research synthesis라는 특정 영역에 집중한다.

다음 요소들은 충분히 측정되지 않는다.

  • organizational politics
  • permissions
  • conflicting stakeholders
  • changing objectives
  • legal accountability
  • multi-month memory
  • high-stakes responsibility
  • long-term ownership

각 benchmark는 deployment에 필요한 한 축을 측정할 뿐이고, 실제 delegability는 여러 축이 동시에 만족될 때 생긴다.

16. 이 글이 주장하지 않는 것

  • 60분 METR horizon이 실제 wall-clock 60분 autonomous operation을 의미한다고 입증하지 않았다.
  • 7개월 doubling trend가 무한히 지속된다고 입증하지 않았다.
  • 50% success가 production-ready threshold라고 입증하지 않았다.
  • GDPval 초기 near-50% 결과가 현재 frontier ceiling이라고 입증하지 않았다.
  • 낮은 inference cost가 낮은 total deployment cost와 동일하다고 입증하지 않았다.
  • DeepScholar-Bench가 영구적으로 19%에 막혀 있다고 입증하지 않았다.
  • software-heavy benchmark 성능이 모든 profession으로 일반화된다고 입증하지 않았다.
  • benchmark 결과만으로 job replacement를 입증하지 않았다.

17. Part 8의 실제 결론

Agent capability가 multidimensional이기 때문에 agent evaluation도 multidimensional이 되고 있다.

어떤 system은 더 긴 task를 수행하지만 자주 실패할 수 있고, professional-looking output을 만들지만 heavy review가 필요할 수 있고, 많은 literature를 검색하지만 citation grounding이 약할 수 있고, low-context benchmark에서는 강하지만 실제 조직 context에서는 약할 수 있다.

그래서 frontier 질문은 “얼마나 똑똑한가?”에서 “얼마나 delegable한가?”로 이동해야 한다.

Delegability에는 최소한 다음이 함께 필요하다.

capability + reliability + context handling + output quality + verifiability + recovery

Agent frontier는 얼마나 오래 일할 수 있는지가 아니다. 실패와 supervision의 기대비용이 autonomy의 가치를 압도하기 전에 얼마나 많은 일을 맡길 수 있는가가 진짜 frontier다.

Part 9에서는 시리즈 마지막인 Future Research Areas로 이동한다.

Claim Map

Speaker / Course Claim

  • long-horizon evaluation은 saturated short benchmark보다 agent capability를 더 잘 드러낸다.
  • task horizon이 길어질수록 reliability가 중요해진다.
  • economic quality와 research verifiability는 short correctness benchmark에서 드러나지 않는 failure를 보여준다.

Verified Fact

  • METR 초기 연구는 50% task-completion horizon의 역사적 doubling time을 약 7개월로 추정했다.
  • Claude 3.7 Sonnet 50% horizon은 setup에 따라 약 55~59분이다.
  • METR time horizon은 agent wall-clock이 아니라 human expert task duration으로 정의된다.
  • METR current methodology는 current task suite에서 16시간 초과 측정을 unreliable하다고 경고한다.
  • GDPval v1은 9 sectors, 44 occupations, 1,320 tasks, 220-task gold subset이며 전문가 평균 경력은 14년 이상이다.
  • GDPval paper는 arXiv:2510.04374이다.
  • later GDPval-family 결과는 초기 near-50% frontier snapshot을 크게 넘어섰다.
  • DeepScholar-Bench는 arXiv:2508.20033이다.
  • early version은 모든 metric에서 19%를 넘은 system이 없었다.
  • later ICLR revision은 geometric mean 31%를 넘은 system이 없었다.

Editorial Interpretation

  • task duration, reliability, economic usefulness, verifiability를 분리해서 측정해야 한다.
  • benchmark success는 controlled dimension에 대한 evidence이지 workforce replacement의 직접 증거가 아니다.
  • agent progress의 실용적 metric은 task length 자체보다 delegability에 가깝다.

강의 지도

  • 00:00–00:28 — METR, time horizon, reliability, low-context limitation
  • 00:28–00:52 — GDPval, professional deliverable quality
  • 00:52–01:08 — DeepScholar-Bench, research synthesis and verifiability
  • 01:08–01:15 — long-horizon reliability and deployment implications

주요 1차 출처