
Scaling Laws
Phillip Isola teaches lecture 20 of MIT's 6.7960 Deep Learning, covering scaling laws in neural network architectures. He walks through power law relationships between model size, dataset size, and compute, then examines where these laws break down. The lecture builds toward theoretical explanations for why scaling works as well as it does, and introduces the concept of critical batch size, the point past which increasing batch size stops yielding proportional training speedups. Aimed at students who already have grounding in deep learning fundamentals, the talk connects empirical observations from large-scale training runs to the underlying math that predicts them. Running 38 minutes, it is part of MIT's Fall 2024 offering of 6.7960 and assumes familiarity with neural network training dynamics covered in earlier sessions of the course.