LECTURES A GRATIS GLOBAL SERVICE
⌕ SEARCH GRATIS GLOBAL ↗
LECTURES
Scaling Laws
SOURCE: YOUTUBE · NO TRACKING UNTIL YOU PRESS PLAY · TROUBLE PLAYING? WATCH AT THE SOURCE ↗

Scaling Laws

38 MIN · EN · STATUS: [ STREAMING ]
RATE THIS
MIT · Deep Learning · LECTURE 20

Phillip Isola teaches lecture 20 of MIT's 6.7960 Deep Learning, covering scaling laws in neural network architectures. He walks through power law relationships between model size, dataset size, and compute, then examines where these laws break down. The lecture builds toward theoretical explanations for why scaling works as well as it does, and introduces the concept of critical batch size, the point past which increasing batch size stops yielding proportional training speedups. Aimed at students who already have grounding in deep learning fundamentals, the talk connects empirical observations from large-scale training runs to the underlying math that predicts them. Running 38 minutes, it is part of MIT's Fall 2024 offering of 6.7960 and assumes familiarity with neural network training dynamics covered in earlier sessions of the course.

More from this course

12 LECTURES
Introduction to Deep Learning

Introduction to Deep Learning

MIT · 61 MIN
How to Train a Neural Net

How to Train a Neural Net

MIT · 80 MIN
Approximation Theory

Approximation Theory

MIT · 83 MIN
Architectures: Grids

Architectures: Grids

MIT · 84 MIN
Architectures: Graphs

Architectures: Graphs

MIT · 81 MIN
Generalization Theory

Generalization Theory

MIT · 81 MIN
Scaling Rules for Optimization

Scaling Rules for Optimization

MIT · 81 MIN
Architectures: Transformers

Architectures: Transformers

MIT · 75 MIN
Hacker's Guide to Deep Learning

Hacker's Guide to Deep Learning

MIT · 76 MIN
Architectures: Memory

Architectures: Memory

MIT · 73 MIN
Lec 11: Representation Learning: Reconstruction-Based

Lec 11: Representation Learning: Reconstruction-Based

MIT · 81 MIN
Representation Learning: Similarity-Based

Representation Learning: Similarity-Based

MIT · 76 MIN