23. What is RLHF? Reinforcement Learning from Human Feedback Explained In Hindi

AI SayI · Advanced ·🎮 Reinforcement Learning ·6mo ago

About this lesson

How do AI models like ChatGPT learn to be so helpful and safe? In this video, we break down Reinforcement Learning from Human Feedback (RLHF)—the essential technique used to align large language models (LLMs) with human values. We dive deep into the technical workflow, covering everything from initial pre-training to advanced policy optimization. What you’ll learn in this video: What RLHF is and why it's better than supervised learning alone. Step 1: Pre-training – The foundation of any LLM. Step 2: Human Feedback Collection – How human ranking and rating shape AI behavior. Step 3: Reward Modeling – Building a system that predicts what humans want. Step 4: Reinforcement Learning – Fine-tuning the model using Policy Optimization (PPO). Step 5: Iteration – The continuous process of improving AI alignment. Whether you are an AI student, a developer, or just curious about how modern LLMs are built, this breakdown will give you a clear understanding of the RLHF pipeline. Don't forget to Like, Subscribe, and hit the Notification Bell for more AI and Machine Learning deep dives!

Original Description

How do AI models like ChatGPT learn to be so helpful and safe? In this video, we break down Reinforcement Learning from Human Feedback (RLHF)—the essential technique used to align large language models (LLMs) with human values. We dive deep into the technical workflow, covering everything from initial pre-training to advanced policy optimization. What you’ll learn in this video: What RLHF is and why it's better than supervised learning alone. Step 1: Pre-training – The foundation of any LLM. Step 2: Human Feedback Collection – How human ranking and rating shape AI behavior. Step 3: Reward Modeling – Building a system that predicts what humans want. Step 4: Reinforcement Learning – Fine-tuning the model using Policy Optimization (PPO). Step 5: Iteration – The continuous process of improving AI alignment. Whether you are an AI student, a developer, or just curious about how modern LLMs are built, this breakdown will give you a clear understanding of the RLHF pipeline. Don't forget to Like, Subscribe, and hit the Notification Bell for more AI and Machine Learning deep dives!
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Learn how to improve off-policy reinforcement learning with auxiliary branches, enhancing reasoning in large language models
ArXiv cs.AI
📰
A Practical Guide to Implementing the REINFORCE Algorithm in Python (Part 5)
Implement the REINFORCE algorithm in Python using PyTorch and Gymnasium for reinforcement learning tasks
Medium · Machine Learning
📰
Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies
Learn how to test reinforcement learning policies with Gimitest, a comprehensive tool for ensuring reliability and safety
ArXiv cs.AI
📰
RLVP: Penalize the Path, Reward the Outcome
Learn how to implement RLVP, a new reinforcement learning approach that prioritizes outcome over path, and apply it to real-world problems with costly interactions
ArXiv cs.AI
Up next
How Netflix Uses Reinforcement Learning to Recommend Movies #ai #coding #machinelearning #netflix
Ascent
Watch →