OpenAI DevDay 2026: AI Is Moving From Chat to Persistent Agents
OpenAI’s DevDay 2026 points to a larger shift: from prompt-response chatbots to persistent agents with durable sessions, cloud execution, software access and lower-cost models.
OpenAI’s DevDay 2026 points to a larger shift: from prompt-response chatbots to persistent agents with durable sessions, cloud execution, software access and lower-cost models.
Stanford CS329A Part 9 examines future research in self-improving AI: reasoning diversity, verification, self-generated curricula, continual learning, and intelligence per watt.
Stanford CS329A Part 8 examines long-horizon agent evaluation through METR, GDPval, and DeepScholar-Bench, arguing that longer task horizons are not the same as reliable, delegable AI.
Stanford CS329A Part 7 connects AlphaCode, AlphaCode 2, Search-o1, and Search-R1 to show why search only helps when generation, selection, and context refinement improve together.
Stanford CS329A Part 6 connects STaR, DeepSeekMath/GRPO, and DAPO to show how verified reasoning can be turned into persistent train-time improvement.
Stanford CS329A Part 5 compares METR, GDPval, and DeepScholarBench to show why agent capability, reliability, context, and real-world deliverable quality must be evaluated separately.
Stanford CS329A Part 4 connects ReAct, execution feedback, and Constitutional AI to one question: where should a self-improving agent get corrective signals it can actually trust?
Stanford CS329A Part 3 traces robust verification from outcome verifiers and process reward models to Math-Shepherd, weak-verifier ensembles, and verifier distillation.
Stanford CS329A Part 2 shows why inference scaling is not just about generating more samples, but about allocating compute across search, verification, revision, fusion, and architecture design.
Stanford CS329A Part 1 traces the shift from pretraining scale to chain-of-thought, post-training, inference-time compute, verifiers, and agentic feedback loops.