Train-Time Scaling: How AI Learns from Its Own Reasoning — Stanford CS329A Part 6
Stanford CS329A Part 6 connects STaR, DeepSeekMath/GRPO, and DAPO to show how verified reasoning can be turned into persistent train-time improvement.
Stanford CS329A Part 6 connects STaR, DeepSeekMath/GRPO, and DAPO to show how verified reasoning can be turned into persistent train-time improvement.
Stanford CS329A Part 5 compares METR, GDPval, and DeepScholarBench to show why agent capability, reliability, context, and real-world deliverable quality must be evaluated separately.
Stanford CS329A Part 4 connects ReAct, execution feedback, and Constitutional AI to one question: where should a self-improving agent get corrective signals it can actually trust?
Stanford CS329A Part 3 traces robust verification from outcome verifiers and process reward models to Math-Shepherd, weak-verifier ensembles, and verifier distillation.
Stanford CS329A Part 2 shows why inference scaling is not just about generating more samples, but about allocating compute across search, verification, revision, fusion, and architecture design.
Stanford CS329A Part 1 traces the shift from pretraining scale to chain-of-thought, post-training, inference-time compute, verifiers, and agentic feedback loops.
Stanford’s CS329A shows why self-improving AI is less about one smarter model than a closed loop of search, verification, tools, learning, and long-horizon evaluation.
In 2022, Elon Musk linked autonomy, humanoid robots, brain-computer interfaces and Starship into one optimistic future. Four years later, their uneven progress shows why vision, technical direction and timelines should be judged separately.
A 2017 TED conversation connects tunnels, electric cars, solar energy, reusable rockets and Mars through one larger idea: progress should solve real problems and make the future feel worth anticipating.
AI data centers are forcing utilities and regulators to solve two problems at once: how to add power quickly, and how to keep the cost and risk of new infrastructure from landing on ordinary ratepayers.