
CS25: Transformers United - Overview of Transformers
Stanford's CS25 seminar opens with a guided history of Transformers, the neural network architecture behind modern NLP and large language models. Instructors Steven Feng and Karan P. Singh, both Stanford PhD students, trace the field from earlier machine learning and NLP approaches through the 2017 attention mechanism breakthrough, with guest remarks from professors Michael C. Frank and Christopher Manning situating the architecture within linguistics and cognitive science. The session covers how self-attention and encoder-decoder structures work mechanically, then moves into current applications across text, vision, and multimodal systems, along with open challenges such as scaling, efficiency, and interpretability. Pitched at a technical audience already familiar with basic deep learning, the talk functions as the orientation lecture for the rest of the CS25 series, mapping out the terrain before later sessions go deeper into specific variants and use cases.