Verifier도 틀릴 수 있다 — Stanford CS329A Part 3 Robust Verification
Stanford CS329A Part 3를 따라 outcome verifier, process reward model, Math-Shepherd, weak-verifier ensemble과 distillation까지 검증 기술의 진화를 살펴본다.
Stanford CS329A Part 3를 따라 outcome verifier, process reward model, Math-Shepherd, weak-verifier ensemble과 distillation까지 검증 기술의 진화를 살펴본다.
Stanford CS329A Part 3 traces robust verification from outcome verifiers and process reward models to Math-Shepherd, weak-verifier ensembles, and verifier distillation.
Stanford CS329A Part 2를 따라 repeated sampling, verifier bottleneck, compute-optimal inference, Archon까지 살펴보면 test-time scaling은 결국 연산 배분 문제라는 점이 드러난다.
Stanford CS329A Part 2 shows why inference scaling is not just about generating more samples, but about allocating compute across search, verification, revision, fusion, and architecture design.
Stanford CS329A Part 1을 따라 사전학습 스케일링에서 CoT, post-training, test-time compute, verifier, agentic feedback loop로 이어지는 변화를 자세히 살펴본다.
Stanford CS329A Part 1 traces the shift from pretraining scale to chain-of-thought, post-training, inference-time compute, verifiers, and agentic feedback loops.
Stanford CS329A의 9편 강의를 관통해 보면 자가 개선 AI의 핵심은 더 큰 모델 하나보다 탐색·검증·도구·학습·평가를 연결하는 닫힌 피드백 루프에 가깝다.
Stanford’s CS329A shows why self-improving AI is less about one smarter model than a closed loop of search, verification, tools, learning, and long-horizon evaluation.
Gita Gopinath의 Dominant Currency Paradigm은 왜 양국 환율 변화만으로 실제 무역가격과 무역수지가 교과서처럼 움직이지 않는지 설명한다.
Gita Gopinath’s dominant currency paradigm explains why a bilateral exchange-rate move may do much less to trade prices than the textbook story suggests.