
Tutorial Session: Review of Q-Learning
Anikait Singh, teaching assistant for Stanford's CS224R Deep Reinforcement Learning course, leads this tutorial session recorded in April 2025. He reviews Markov Decision Processes before moving into exact MDP solutions and parametric Q-learning, walking through the math step by step on the board. The session covers practical details students need for implementing these algorithms, then works through a full Q-learning and actor-critic algorithm walkthrough, tying the theory to code-level considerations. This is a supplementary session meant to reinforce lecture material rather than introduce brand new concepts, so it moves briskly through foundational reinforcement learning ideas before settling into the harder parametric and actor-critic material. The format is a standard whiteboard-style tutorial aimed at students already following the course, with Singh fielding the kind of implementation questions that come up when moving from exact tabular methods to function approximation.