Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 7 - Evaluation
Skills:
Modern CV Models90%
Key Takeaways
Explores evaluation methods for diffusion and large vision models in the context of image AI
Original Description
Learn more details about this course: https://online.stanford.edu/courses/cme296-diffusion-and-large-vision-models
To follow along with the course schedule and syllabus, visit: https://cme296.stanford.edu/syllabus/
Chapters:
00:00:00 Introduction
00:05:19 Motivation
00:10:48 Human ratings
00:19:43 Elo rating system
00:26:37 Reference-free metrics
00:29:15 Fréchet inception distance (FID)
00:42:30 CLIPScore
00:44:51 PickScore
00:45:41 Reference-based metrics
00:48:07 Mean squared error (MSE)
00:49:36 Peak signal-to-noise ratio (PSNR)
00:51:54 Structural similarity (SSIM)
01:01:09 Perceptual similarity (LPIPS)
01:05:03 Multimodal LLMs
01:13:10 Faithfulness evaluation (TIFA)
01:17:29 Visual question answering score (VQA)
01:24:40 MLLM-as-a-Judge
01:34:17 Benchmarks
For more information about Stanford’s graduate programs, visit: https://online.stanford.edu/graduate-education
Afshine Amidi is an Adjunct Lecturer at Stanford University.
Shervine Amidi is an Adjunct Lecturer at Stanford University.
View the course playlist: https://www.youtube.com/playlist?list=PLoROMvodv4rNdy8rt2rZ4T2xM0OjADnfu
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: Modern CV Models
View skill →Related Reads
Chapters (18)
Introduction
5:19
Motivation
10:48
Human ratings
19:43
Elo rating system
26:37
Reference-free metrics
29:15
Fréchet inception distance (FID)
42:30
CLIPScore
44:51
PickScore
45:41
Reference-based metrics
48:07
Mean squared error (MSE)
49:36
Peak signal-to-noise ratio (PSNR)
51:54
Structural similarity (SSIM)
1:01:09
Perceptual similarity (LPIPS)
1:05:03
Multimodal LLMs
1:13:10
Faithfulness evaluation (TIFA)
1:17:29
Visual question answering score (VQA)
1:24:40
MLLM-as-a-Judge
1:34:17
Benchmarks
🎓
Tutor Explanation
DeepCamp AI