
Generative AI: Text-to-Image Models
Rama Ramakrishnan, teaching MIT's 15.773 Hands-On Deep Learning, lectures on text-conditional diffusion models built on transformer architectures. The session covers how these systems turn written prompts into images and video, walking through the mechanics of diffusion based generation and the role transformers play in conditioning outputs on text input. Ramakrishnan works through text-to-image and text-to-video examples to show how the models interpret language and translate it into visual content, situating the techniques within the broader deep learning toolkit covered in the course. The lecture runs seventy six minutes and assumes familiarity with earlier sessions in the course, building on prior material on neural network architectures. It is aimed at students wanting a technical grounding in how modern generative AI image and video systems actually work under the hood, rather than a product demo or high level overview.