Everyone uses AI; few profit from it yet. Agents — systems where the model owns the control flow — are the bet the industry is making to close that gap. This week builds your mental model: the loop, the vocabulary, the autonomy dial, and the 2026 landscape.
Click each card for the “so what” and the source. All figures are approximate and dated — that’s a feature of this field, not a bug of these slides.
The length of tasks agents can complete autonomously has been doubling roughly every three months (METR, approx.). Demos get dramatically better every semester.
On long computer-use tasks, top agents score around 20% where humans score ~72% (OSWorld, approx.). Benchmarks lag the demos — reliability is the semester’s recurring villain.