Same Answer, Different Representations: Hidden instability in VLMs

📰 ArXiv cs.AI

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processing. In this work, we argue that this assumption is insufficient. We introduce a representation-aware and frequency-aware evaluation framework that measures internal embedding drift, spectral sensitivity, and structural smoothness (spatial consistency of vision tokens)

Published 16 Sept 2026
Read full paper → ☆ Save to playlist ← Back to Reads

Related Videos

YOLO V2 | Object Detection Series | Part 2
YOLO V2 | Object Detection Series | Part 2
AGI Lambda
This System Captures The Whole Stadium At Once | AI Chooses The Perfect Shot
This System Captures The Whole Stadium At Once | AI Chooses The Perfect Shot
Anik Singal
Build a WhatsApp AI Agent (Auto Replies) Using OpenClaw – Step-by-Step
Build a WhatsApp AI Agent (Auto Replies) Using OpenClaw – Step-by-Step
Muhammad Moin
Choosing Your Path: AI Professional Program Course Selection Guide
Choosing Your Path: AI Professional Program Course Selection Guide
Stanford Online
TensorFlow: Advanced Techniques Specialization
TensorFlow: Advanced Techniques Specialization
DeepLearning.AI
Multimodal Data Analysis with AI
Multimodal Data Analysis with AI
Latent Space