오래 일할 수 있는 AI와 믿고 맡길 수 있는 AI는 다르다 — Stanford CS329A Part 8
Stanford CS329A Part 8을 따라 METR, GDPval, DeepScholar-Bench를 비교하며 긴 task horizon과 실제로 믿고 맡길 수 있는 AI의 차이를 설명한다.
Stanford CS329A Part 8을 따라 METR, GDPval, DeepScholar-Bench를 비교하며 긴 task horizon과 실제로 믿고 맡길 수 있는 AI의 차이를 설명한다.
Stanford CS329A Part 8 examines long-horizon agent evaluation through METR, GDPval, and DeepScholar-Bench, arguing that longer task horizons are not the same as reliable, delegable AI.
Stanford CS329A Part 7을 따라 AlphaCode, AlphaCode 2, Search-o1, Search-R1을 연결한다. Search의 핵심은 더 많이 찾는 것이 아니라 더 잘 생성·선별·압축하는 데 있다.
Stanford CS329A Part 7 connects AlphaCode, AlphaCode 2, Search-o1, and Search-R1 to show why search only helps when generation, selection, and context refinement improve together.
Stanford CS329A Part 6를 따라 STaR, DeepSeekMath/GRPO, DAPO를 연결한다. 검증된 reasoning이 어떻게 지속적인 train-time improvement로 바뀌는지 살펴본다.
Stanford CS329A Part 6 connects STaR, DeepSeekMath/GRPO, and DAPO to show how verified reasoning can be turned into persistent train-time improvement.
Stanford CS329A Part 5의 METR, GDPval, DeepScholarBench를 통해 agent capability와 reliability, context, 전문가 수준 산출물 품질이 왜 별개인지 살펴본다.
Stanford CS329A Part 5 compares METR, GDPval, and DeepScholarBench to show why agent capability, reliability, context, and real-world deliverable quality must be evaluated separately.
Stanford CS329A Part 4를 따라 ReAct, 코드 실행 피드백, Constitutional AI를 하나의 질문으로 연결한다. Self-improving agent는 어떤 교정 신호를 믿어야 하는가?
Stanford CS329A Part 4 connects ReAct, execution feedback, and Constitutional AI to one question: where should a self-improving agent get corrective signals it can actually trust?