Case study on CLIP: Large Multi-Modal Models for Blind & Low Vision Users | Microsoft Research Forum

Microsoft Research · Advanced ·👁️ Computer Vision ·2y ago

Key Takeaways

The video discusses the challenges and opportunities of large multi-modal models, such as CLIP, for blind and low vision users, highlighting the performance disparities of these models on data captured by blind users.

Full Transcript

[Music] hi there my name is Daniela masetti and I'm a senior researcher at Microsoft research Cambridge today I will be sharing our recent cvpr paper which examines the challenges and opportunities of large multile models for blind and low vision users today's AI models hold incredible potential for assisting the blind Community from text recognition to object identification to question answering acts like seeing AI are already deploying some of these AI features but there is potential for much more and I think this is hinted at by the recent partnership between openai and bmis with the promise that one day Human Assistance could be replaced by AI agents that provide instantaneous assistance to Blind users around the world but despite their potential no Works have really looked at well how well do these models actually work on image and Text data captured by blind users and we know from the literature that this data is likely to be out of distribution or different in a number of ways for example blind users use a range of quite specialized assistive objects they also are more likely to capture images with quality variation things like Camera blur and occlusion and they're also more likely to make use of non-visual vocabulary for example describing their objects by their physical rather than their visual properties our work therefore set out to remedy this specifically we systematically evaluated 25 variants of the clip model on data data from blind and low vision users clip is one of today's most widely used multimodal models it has over 15,000 citations and 75 million downloads we used the orbit and the VIS classification data sets both of these are collected by blind users through real world assistive applications and we inspected Clips performance on both a zero shot image classification task directly as well as through examining the performance of models that use clip as a component which is very widely done in the community I unfortunately don't have time to go into all the details of our work uh but I will share our top three findings with you first we confirmed that clip does indeed underperform on data that is captured by blind and low vision users second these disparities trickle down to models that use clip as a component and then third these disparities stem from the fact that disability content is significantly under represented and some sometimes missing completely from the data sets that are used to pre-train these large models I'll now dive into our three findings in a bit more detail so for the first finding we found that clip underperforms on objects image quality and language that is typically used by blind users on object type clip recognizes disability objects like a braille keyboard for example up to 28 percentage points less accurately than common objects like a TV remote on image quality clip is up to 23 percentage points more sensitive to images with things like Camera blur and lighting compared to images that don't have these quality issues and on language clip recognizes objects that are described by their material so for example a leather boot up to 12 percentage points less accurately than objects described by their color for example a brown Boot and we know that blind users rely heavily on this tactile rather than visual language towards our second finding We examined three models that use clip under the hood an object detection model an image segmentation model and an image generation model and found that all three struggle with disability content for example doly 2 which relies on a clip Vision encoder cannot generate common disability objects like guide canes and Braille keyboards instead as you can see here it gives us very strange looking walking canes and lots and lots of randomly placed white dots in comparison delu generated really high quality and realistic images for almost all of the non-disability objects that we tested and then towards our third and final finding we really wanted to understand where these performance disparities were stemming from and so we Quantified just how prevalent disability content is in three popular data sets that are commonly used to pre-train these large models lion 4 million lion 2 billion and the data comp 1B data set or 1 billion data set specifically we counted how many times objects are mentioned in these data sets captions and found that disability objects appear up to 16 to 17 times less frequently than non-disability objects across all three of the data sets so as you can see our work has identified a clear Gap in current models capabilities for blind users and this could have very real consequences if these models are then integrated into assistive Technologies for the Blind and low vision community so what should we as a research community be doing about it first I think more work is needed to understand how models come to learn or adapt to longtailed data some of our early results show that F shot learning approach approaches hold some promise but they don't always work especially in more challenging scenarios for example when objects appear in highly cluttered scenarios and second I think it's important for us to really focus on including more disability content in these large scale pre-training data sets and our team are currently working on developing Equitable and fair practices alongside disabled communities to Source data that is truly representative of their needs and so with that I will wrap up thank you to all the people behind this work and thank you for listening

Original Description

Insights into the Challenges and Opportunities of Large Multi-Modal Models for Blind and Low Vision Users: A Case Study on CLIP Daniela Massiceti delves into the transformative potential of multimodal models such as CLIP for assistive technologies. Specifically focusing on the blind/low-vision community, the talk explores the current distance from realizing this potential and the advancements needed to bridge this gap. This session aired on June 4, 2024 at Microsoft Research Forum, Episode 3. Register for the series: https://aka.ms/registerresearchforumYTe3 Continue watching episode 3: https://aka.ms/researchforumYTe3 Explore all previous episodes: https://aka.ms/researchforumYTplaylist
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Microsoft Research · Microsoft Research · 0 of 60

← Previous Next →
1 Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Microsoft Research
2 Frontiers in Machine Learning: Climate Impact of Machine Learning
Frontiers in Machine Learning: Climate Impact of Machine Learning
Microsoft Research
3 Frontiers in Machine Learning: Security and Machine Learning
Frontiers in Machine Learning: Security and Machine Learning
Microsoft Research
4 Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Microsoft Research
5 Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Microsoft Research
6 Remote Work and Well-Being
Remote Work and Well-Being
Microsoft Research
7 Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Microsoft Research
8 Towards a Practical Virtual Office for Mobile Knowledge Workers
Towards a Practical Virtual Office for Mobile Knowledge Workers
Microsoft Research
9 Impact of COVID-19 crisis on the future of work in India
Impact of COVID-19 crisis on the future of work in India
Microsoft Research
10 Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Microsoft Research
11 How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
Microsoft Research
12 Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Microsoft Research
13 Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Microsoft Research
14 Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Research
15 Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Research
16 Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Research
17 Directions in ML: Algorithmic foundations of neural architecture search
Directions in ML: Algorithmic foundations of neural architecture search
Microsoft Research
18 MineRL Competition 2020
MineRL Competition 2020
Microsoft Research
19 Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Microsoft Research
20 From Paper to Product
From Paper to Product
Microsoft Research
21 SkinnerDB: Regret Bounded Query Evaluation using RL
SkinnerDB: Regret Bounded Query Evaluation using RL
Microsoft Research
22 From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
Microsoft Research
23 Programming with Proofs for High-assurance Software
Programming with Proofs for High-assurance Software
Microsoft Research
24 Platform for Situated Intelligence Overview
Platform for Situated Intelligence Overview
Microsoft Research
25 Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Microsoft Research
26 Galactic Bell Star Music Demo
Galactic Bell Star Music Demo
Microsoft Research
27 Importing Animations in Microsoft Expressive Pixels (9 of 9)
Importing Animations in Microsoft Expressive Pixels (9 of 9)
Microsoft Research
28 Welcome to Microsoft Expressive Pixels (1 of 9)
Welcome to Microsoft Expressive Pixels (1 of 9)
Microsoft Research
29 Getting Started with Microsoft Expressive Pixels (2 of 9)
Getting Started with Microsoft Expressive Pixels (2 of 9)
Microsoft Research
30 Creating an Image in Microsoft Expressive Pixels (3 of 9)
Creating an Image in Microsoft Expressive Pixels (3 of 9)
Microsoft Research
31 Creating Animations in Microsoft Expressive Pixels (4 of 9)
Creating Animations in Microsoft Expressive Pixels (4 of 9)
Microsoft Research
32 Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Microsoft Research
33 Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Microsoft Research
34 Using Layers in Microsoft Expressive Pixels (7 of 9)
Using Layers in Microsoft Expressive Pixels (7 of 9)
Microsoft Research
35 Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Microsoft Research
36 What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
Microsoft Research
37 What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
Microsoft Research
38 Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Microsoft Research
39 Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Microsoft Research
40 Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Microsoft Research
41 Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Microsoft Research
42 Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Microsoft Research
43 Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Microsoft Research
44 Novel Image Captioning
Novel Image Captioning
Microsoft Research
45 Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Microsoft Research
46 Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Microsoft Research
47 How does holographic storage work?
How does holographic storage work?
Microsoft Research
48 The physics of hologram formation in iron doped lithium niobate
The physics of hologram formation in iron doped lithium niobate
Microsoft Research
49 Introduction to coax: A Modular RL Package
Introduction to coax: A Modular RL Package
Microsoft Research
50 Directions in ML: "Neural architecture search: Coming of age"
Directions in ML: "Neural architecture search: Coming of age"
Microsoft Research
51 Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research
52 Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Microsoft Research
53 Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Microsoft Research
54 Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research
55 Clinical Research with FHIR
Clinical Research with FHIR
Microsoft Research
56 Soundscape Street Preview
Soundscape Street Preview
Microsoft Research
57 Tilt-Responsive Techniques for Digital Drawing Boards
Tilt-Responsive Techniques for Digital Drawing Boards
Microsoft Research
58 SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
Microsoft Research
59 Haptic PIVOT: On-Demand Handhelds in VR
Haptic PIVOT: On-Demand Handhelds in VR
Microsoft Research
60 SurfaceFleet Supplemental Video Demonstration (UIST 2020)
SurfaceFleet Supplemental Video Demonstration (UIST 2020)
Microsoft Research

This video discusses the challenges and opportunities of large multi-modal models for blind and low vision users, highlighting the need for more inclusive and equitable AI practices. The speaker presents a case study on CLIP, a widely used multimodal model, and its performance on data captured by blind users.

Key Takeaways
  1. Collect and analyze data from blind and low vision users
  2. Evaluate the performance of multimodal models on this data
  3. Develop and implement more inclusive and equitable AI practices
  4. Source data that is truly representative of the needs of blind and low vision users
  5. Use few-shot learning approaches to adapt to long-tailed data
💡 The performance disparities of large multi-modal models on data captured by blind users are significant, and more work is needed to develop inclusive and equitable AI practices.

Related Reads

Up next
9-Phase Computer Vision Roadmap 2026 | AI & Deep Learning | #shorts
SCALER
Watch →