All
Articles 132,997Blog Posts 137,553Tech Tutorials 34,494Research Papers 25,940News 18,844
⚡ AI Lessons
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
1y ago
Why We Think
Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute ( Graves et al. 2016 , Ling, et al. 2017 ,
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
1y ago
Reward Hacking in Reinforcement Learning
Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely l
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
2y ago
Extrinsic Hallucinations in LLMs
Hallucination in large language models usually refers to the model generating unfaithful, fabricated, inconsistent, or nonsensical content. As a term, hallucina
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
2y ago
Adversarial Attacks on LLMs
The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have in
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
3y ago
LLM Powered Autonomous Agents
Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT , GPT-Engineer and Ba
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
3y ago
Prompt Engineering
Prompt Engineering , also known as In-Context Prompting , refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without u
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
3y ago
The Transformer Family Version 2.0
Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
3y ago
Large Transformer Model Inference Optimization
[Updated on 2023-01-24: add a small section on Distillation .] Large transformer models are mainstream nowadays, creating SoTA results for a variety of tasks. T
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
3y ago
Some Math behind Neural Tangent Kernel
Neural networks are well known to be over-parameterized and can often easily fit data with near-zero training loss with decent generalization performance on tes
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
4y ago
Learning with not Enough Data Part 2: Active Learning
This is part 2 of what to do when facing a limited amount of labeled data for supervised learning tasks. This time we will get some amount of human labeling wor
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
4y ago
How to Train Really Large Models on Many GPUs?
[Updated on 2022-03-13: add expert choice routing .] [U
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
5y ago
Contrastive Representation Learning
The goal of contrastive representation learning is to learn such an embedding space in which similar sample pairs stay close to each other while dissimilar ones
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
5y ago
Reducing Toxicity in Language Models
Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
6y ago
Exploration Strategies in Deep Reinforcement Learning
[Updated on 2020-06-17: Add “exploration via disagreement” in the “Forward Dynamics” section . Exploitation versus ex
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
6y ago
Self-Supervised Representation Learning
[Updated on 2020-01-09: add a new section on Contrastive Predictive Coding ]. [Updated on 2020-04-13: add a “Momentum Contra
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
6y ago
Evolution Strategies
Stochastic gradient descent is a universal choice for optimizing deep learning models. However, it is not the only option. With bl
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
7y ago
Meta Reinforcement Learning
In my earlier post on meta-learning , the problem is mainly defined in the context of few-shot classificati
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
7y ago
Domain Randomization for Sim2Real Transfer
In Robotics, one of the hardest problems is how to make your model transfer to the real world. Due to the sample inefficiency of deep RL algorithms and the cost
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
7y ago
Generalized Language Models
[Updated on 2019-02-14: add ULMFiT and GPT-2 .] [Updated on 2020-02-29: add ALBERT .] <span class="updat
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
7y ago
Meta-Learning: Learning to Learn Fast
[Updated on 2019-10-01: thanks to Tianhao, we have
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
7y ago
Flow-based Deep Generative Models
So far, I’ve written about two types of generative models, GAN and VAE . Neither of them explicitly learns the probability density function of real da
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
8y ago
Implementing Deep Reinforcement Learning Models with Tensorflow + OpenAI Gym
The full implementation is available in lilianweng/deep-reinforcement-learning-gym In the previous two posts, I have introduced the algorithms of many deep rein
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
8y ago
Policy Gradient Algorithms
[Updated on 2018-06-30: add two new policy gradient methods, SAC and D4PG .] [Updated on 2018-09-30: ad
Lilian Weng's Blog
🧠 Large Language Models
⚡ AI Lesson
8y ago
A (Long) Peek into Reinforcement Learning
[Updated on 2020-09-03: Updated the algorithm of SARSA and Q-learning so that the diffe
DeepCamp AI