✕ Clear filters
714 lessons

🎮 Reinforcement Learning

RL algorithms, reward modelling, RLHF, policy gradients, Q-learning and multi-agent RL

All ▶ YouTube 331,609📚 External: Coursera 20,471🏛 Archive.org 649 | 📰 Articles →

Looking for written articles and micro-lessons? Switch to Reads.

Middle Management Meritocracy: Shockingly Naive
Reinforcement Learning ⚡ AI Lesson
Middle Management Meritocracy: Shockingly Naive
iBankerU Intermediate 2mo ago
Off-Leash Reliability: A 10-Minute Guide to Real Trust
Reinforcement Learning ⚡ AI Lesson
Off-Leash Reliability: A 10-Minute Guide to Real Trust
UBC News Business Beginner 2mo ago
AI Agents: The Definitive Guide — Chapter 3: Advanced RL & Sequence Learning
AI Agents & Automation ⚡ AI Lesson
AI Agents: The Definitive Guide — Chapter 3: Advanced RL & Sequence Learning
onepagecode Advanced 2mo ago
Reinforcement Learning : Agent, Environment, Action, Reward, Policy Simply Explained
ML Fundamentals ⚡ AI Lesson
Reinforcement Learning : Agent, Environment, Action, Reward, Policy Simply Explained
codehubgenius Beginner 2mo ago
Reinforcement Learning Tutorial Part 9
Reinforcement Learning ⚡ AI Lesson
Reinforcement Learning Tutorial Part 9
Stephen Blum Beginner 1mo ago
PPO - Why Big Steps Break Reinforcement Learning
ML Fundamentals ⚡ AI Lesson
PPO - Why Big Steps Break Reinforcement Learning
DataMListic Beginner 1mo ago
What is RLHF (Reinforcement Learning from Human Feedback) ? | The Secret Ingredient Behind ChatGPT
2:15
Reinforcement Learning ⚡ AI Lesson
What is RLHF (Reinforcement Learning from Human Feedback) ? | The Secret Ingredient Behind ChatGPT
VLR Software Training Beginner 10mo ago
What is Reinforcement Learning from Human Feedback (RLHF)
0:54
Reinforcement Learning ⚡ AI Lesson
What is Reinforcement Learning from Human Feedback (RLHF)
Data Science Made Easy Beginner 10mo ago
Training a Unitree G1 to Walk w/ Reinforcement Learning
Reinforcement Learning ⚡ AI Lesson
Training a Unitree G1 to Walk w/ Reinforcement Learning
Sentdex Advanced 9mo ago
Why Most People Stay Poor
Reinforcement Learning ⚡ AI Lesson
Why Most People Stay Poor
Dan Lok Intermediate 6mo ago
The Hidden Power Behind Loyalty Programs 🤔    #shorts
Reinforcement Learning ⚡ AI Lesson
The Hidden Power Behind Loyalty Programs 🤔 #shorts
Jacky Chou from Indexsy Beginner 8mo ago
Neuro-symbolic AI: Reservoir computing + Reinforcement Learning | Hands-on
Reinforcement Learning ⚡ AI Lesson
Neuro-symbolic AI: Reservoir computing + Reinforcement Learning | Hands-on
BrainOmega Beginner 7mo ago
Supervised vs Unsupervised vs Reinforcement Learning
ML Fundamentals ⚡ AI Lesson
Supervised vs Unsupervised vs Reinforcement Learning
Analytics Vidhya Beginner 5mo ago
Learn with Me: Train AI Agents for Command-Line Tasks with Synthetic Data and RL | Nemotron Labs
Reinforcement Learning ⚡ AI Lesson
Learn with Me: Train AI Agents for Command-Line Tasks with Synthetic Data and RL | Nemotron Labs
NVIDIA Developer Beginner 8mo ago
How ChatGPT Actually Works: The "Secret Sauce" of AI Alignment & RLHF Explained
Reinforcement Learning ⚡ AI Lesson
How ChatGPT Actually Works: The "Secret Sauce" of AI Alignment & RLHF Explained
The Latent Space Beginner 10mo ago
RLHF: Reinforcement Learning from Human Feedback - An explainer for Humans - AI Tasks/Annotators
6:01
Reinforcement Learning ⚡ AI Lesson
RLHF: Reinforcement Learning from Human Feedback - An explainer for Humans - AI Tasks/Annotators
TheCatWith7Legs Beginner 9mo ago
LLM Fine-Tuning 17: Fine-Tune ANY LLM with LLaMA Factory | Full Guide (WebUI + CLI | LoRA + QLoRA)
Reinforcement Learning ⚡ AI Lesson
LLM Fine-Tuning 17: Fine-Tune ANY LLM with LLaMA Factory | Full Guide (WebUI + CLI | LoRA + QLoRA)
Sunny Savita Beginner 9mo ago
Money Lessons From Americans Caring for Aging Parents | Life Lessons | Business Insider
Reinforcement Learning ⚡ AI Lesson
Money Lessons From Americans Caring for Aging Parents | Life Lessons | Business Insider
Business Insider Beginner 10mo ago
Toblerone’s Hidden Bear Secret 🐻    #shorts
Reinforcement Learning ⚡ AI Lesson
Toblerone’s Hidden Bear Secret 🐻 #shorts
Jacky Chou from Indexsy Beginner 6mo ago
Intelligent Robots in 2026: Are We There Yet? [Nikita Rudin] - 760
Reinforcement Learning ⚡ AI Lesson
Intelligent Robots in 2026: Are We There Yet? [Nikita Rudin] - 760
TWIML AI Podcast Beginner 8mo ago
Train a Reasoning Model for $1.23 (Reinforcement Learning)
Reinforcement Learning ⚡ AI Lesson
Train a Reasoning Model for $1.23 (Reinforcement Learning)
Shane | LLM Implementation Intermediate 8mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 6: Q-Learning
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 6: Q-Learning
Stanford Online Beginner 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 17: Advancing Robot Intelligence
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 17: Advancing Robot Intelligence
Stanford Online Beginner 9mo ago
LLM Fine-Tuning Crash Course: Finetune model on PDFs, Instruction FT, Preference Training (DPO/RLHF)
3:36:14
Reinforcement Learning ⚡ AI Lesson
LLM Fine-Tuning Crash Course: Finetune model on PDFs, Instruction FT, Preference Training (DPO/RLHF)
Sunny Savita Advanced 9mo ago
The Top 1% Think Like THIS | Here's How To Do It | Denis Waitley
Reinforcement Learning ⚡ AI Lesson
The Top 1% Think Like THIS | Here's How To Do It | Denis Waitley
Evan Carmichael Beginner 6mo ago
#Meta rolls out new performance program with bigger #rewards for top #employees
Reinforcement Learning ⚡ AI Lesson
#Meta rolls out new performance program with bigger #rewards for top #employees
Business Insider Advanced 8mo ago
How to train Multi Agent Collaborative Agents with Reinforcement Learning (CTDE Explained)
Reinforcement Learning ⚡ AI Lesson
How to train Multi Agent Collaborative Agents with Reinforcement Learning (CTDE Explained)
Neural Breakdown with AVB Beginner 9mo ago
Direct Preference Optimization (DPO): End-to-End Implementation
Reinforcement Learning ⚡ AI Lesson
Direct Preference Optimization (DPO): End-to-End Implementation
SH AI Academy Intermediate 3mo ago
This AI Learns in Two Minds (Slow RL, Fast GEPA)
Research Papers Explained ⚡ AI Lesson
This AI Learns in Two Minds (Slow RL, Fast GEPA)
Discover AI Beginner 4mo ago
Next Stage of AI Scientist: NanoResearch (Skills, Mem, RL)
Research Papers Explained ⚡ AI Lesson
Next Stage of AI Scientist: NanoResearch (Skills, Mem, RL)
Discover AI Advanced 4mo ago
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Reinforcement Learning ⚡ AI Lesson
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Latent Space Beginner 7mo ago
The Coloring Book Trend Secretly Teaching Critical Thinking in Kids
Reinforcement Learning ⚡ AI Lesson
The Coloring Book Trend Secretly Teaching Critical Thinking in Kids
UBC News Business Beginner 3mo ago
Reinforcement Learning: A (practical) introduction
Reinforcement Learning ⚡ AI Lesson
Reinforcement Learning: A (practical) introduction
Shawhin Talebi Beginner 8mo ago
Why Every Skyrim AI Becomes a Stealth Archer
Reinforcement Learning ⚡ AI Lesson
Why Every Skyrim AI Becomes a Stealth Archer
Siraj Raval Advanced 9mo ago
Rewarding Hard Work and Value Creation
Reinforcement Learning ⚡ AI Lesson
Rewarding Hard Work and Value Creation
Dan Martell Intermediate 3mo ago
Understanding Reinforcement Learning with Prime Intellect and Unsloth | Nemotron Labs
Reinforcement Learning ⚡ AI Lesson
Understanding Reinforcement Learning with Prime Intellect and Unsloth | Nemotron Labs
NVIDIA Developer Advanced 5mo ago
“You Don’t Care About Your Health” #wait
Reinforcement Learning ⚡ AI Lesson
“You Don’t Care About Your Health” #wait
Dr Sermed Mezher Intermediate 3mo ago
Reward Modeling: How to Train a Reward Model for LLMs
Reinforcement Learning ⚡ AI Lesson
Reward Modeling: How to Train a Reward Model for LLMs
SH AI Academy Intermediate 3mo ago
The future-ready employee is not waiting for permission.
Reinforcement Learning ⚡ AI Lesson
The future-ready employee is not waiting for permission.
Future Ready Leadership With Jacob Morgan Beginner 3mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 5: Off-Policy Actor Critic
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 5: Off-Policy Actor Critic
Stanford Online Beginner 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Tutorial Session: Review of Q-Learning
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Tutorial Session: Review of Q-Learning
Stanford Online Beginner 9mo ago
Smarter AI Gradients: How Agents Learn to Think
Reinforcement Learning ⚡ AI Lesson
Smarter AI Gradients: How Agents Learn to Think
Discover AI Beginner 7mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 8: Reward Learning
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 8: Reward Learning
Stanford Online Beginner 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 15: Hierarchical RL and IL
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 15: Hierarchical RL and IL
Stanford Online Intermediate 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 4: Actor-Critic Methods
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 4: Actor-Critic Methods
Stanford Online Intermediate 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 1: Class Intro
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 1: Class Intro
Stanford Online Beginner 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 3: Policy Gradients
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 3: Policy Gradients
Stanford Online Intermediate 9mo ago
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 13: Meta RL
Reinforcement Learning ⚡ AI Lesson
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 13: Meta RL
Stanford Online Intermediate 9mo ago