
Deep Learning for Natural Language: Transformers, Self-Supervised Learning
Rama Ramakrishnan continues MIT's 15.773 Hands-On Deep Learning course with a lecture focused on transformer architectures and self-supervised learning for natural language tasks. Building on earlier sessions, Ramakrishnan works through how transformers process sequential text data, the attention mechanisms that let them model relationships between words, and why self-supervised pretraining on large unlabeled corpora has become the dominant strategy for building language models. The seventy-seven minute session is pitched at a practical, hands-on level consistent with the course's applied focus, aimed at students who already have grounding in neural networks and want to understand the mechanics behind modern NLP systems. Expect whiteboard or slide-based explanation of model internals rather than a conceptual survey, tying the architecture back to how these models get trained and used in practice. Part of MIT OpenCourseWare's Spring 2024 offering of 15.773, taught by Rama Ramakrishnan.