Case study on CLIP: Large Multi-Modal Models for Blind & Low Vision Users | Microsoft Research Forum
Key Takeaways
The video discusses the challenges and opportunities of large multi-modal models, such as CLIP, for blind and low vision users, highlighting the performance disparities of these models on data captured by blind users.
Full Transcript
[Music] hi there my name is Daniela masetti and I'm a senior researcher at Microsoft research Cambridge today I will be sharing our recent cvpr paper which examines the challenges and opportunities of large multile models for blind and low vision users today's AI models hold incredible potential for assisting the blind Community from text recognition to object identification to question answering acts like seeing AI are already deploying some of these AI features but there is potential for much more and I think this is hinted at by the recent partnership between openai and bmis with the promise that one day Human Assistance could be replaced by AI agents that provide instantaneous assistance to Blind users around the world but despite their potential no Works have really looked at well how well do these models actually work on image and Text data captured by blind users and we know from the literature that this data is likely to be out of distribution or different in a number of ways for example blind users use a range of quite specialized assistive objects they also are more likely to capture images with quality variation things like Camera blur and occlusion and they're also more likely to make use of non-visual vocabulary for example describing their objects by their physical rather than their visual properties our work therefore set out to remedy this specifically we systematically evaluated 25 variants of the clip model on data data from blind and low vision users clip is one of today's most widely used multimodal models it has over 15,000 citations and 75 million downloads we used the orbit and the VIS classification data sets both of these are collected by blind users through real world assistive applications and we inspected Clips performance on both a zero shot image classification task directly as well as through examining the performance of models that use clip as a component which is very widely done in the community I unfortunately don't have time to go into all the details of our work uh but I will share our top three findings with you first we confirmed that clip does indeed underperform on data that is captured by blind and low vision users second these disparities trickle down to models that use clip as a component and then third these disparities stem from the fact that disability content is significantly under represented and some sometimes missing completely from the data sets that are used to pre-train these large models I'll now dive into our three findings in a bit more detail so for the first finding we found that clip underperforms on objects image quality and language that is typically used by blind users on object type clip recognizes disability objects like a braille keyboard for example up to 28 percentage points less accurately than common objects like a TV remote on image quality clip is up to 23 percentage points more sensitive to images with things like Camera blur and lighting compared to images that don't have these quality issues and on language clip recognizes objects that are described by their material so for example a leather boot up to 12 percentage points less accurately than objects described by their color for example a brown Boot and we know that blind users rely heavily on this tactile rather than visual language towards our second finding We examined three models that use clip under the hood an object detection model an image segmentation model and an image generation model and found that all three struggle with disability content for example doly 2 which relies on a clip Vision encoder cannot generate common disability objects like guide canes and Braille keyboards instead as you can see here it gives us very strange looking walking canes and lots and lots of randomly placed white dots in comparison delu generated really high quality and realistic images for almost all of the non-disability objects that we tested and then towards our third and final finding we really wanted to understand where these performance disparities were stemming from and so we Quantified just how prevalent disability content is in three popular data sets that are commonly used to pre-train these large models lion 4 million lion 2 billion and the data comp 1B data set or 1 billion data set specifically we counted how many times objects are mentioned in these data sets captions and found that disability objects appear up to 16 to 17 times less frequently than non-disability objects across all three of the data sets so as you can see our work has identified a clear Gap in current models capabilities for blind users and this could have very real consequences if these models are then integrated into assistive Technologies for the Blind and low vision community so what should we as a research community be doing about it first I think more work is needed to understand how models come to learn or adapt to longtailed data some of our early results show that F shot learning approach approaches hold some promise but they don't always work especially in more challenging scenarios for example when objects appear in highly cluttered scenarios and second I think it's important for us to really focus on including more disability content in these large scale pre-training data sets and our team are currently working on developing Equitable and fair practices alongside disabled communities to Source data that is truly representative of their needs and so with that I will wrap up thank you to all the people behind this work and thank you for listening
Original Description
Insights into the Challenges and Opportunities of Large Multi-Modal Models for Blind and Low Vision Users: A Case Study on CLIP
Daniela Massiceti delves into the transformative potential of multimodal models such as CLIP for assistive technologies. Specifically focusing on the blind/low-vision community, the talk explores the current distance from realizing this potential and the advancements needed to bridge this gap.
This session aired on June 4, 2024 at Microsoft Research Forum, Episode 3.
Register for the series: https://aka.ms/registerresearchforumYTe3
Continue watching episode 3: https://aka.ms/researchforumYTe3
Explore all previous episodes: https://aka.ms/researchforumYTplaylist
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Microsoft Research · Microsoft Research · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Microsoft Research
Frontiers in Machine Learning: Climate Impact of Machine Learning
Microsoft Research
Frontiers in Machine Learning: Security and Machine Learning
Microsoft Research
Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Microsoft Research
Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Microsoft Research
Remote Work and Well-Being
Microsoft Research
Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Microsoft Research
Towards a Practical Virtual Office for Mobile Knowledge Workers
Microsoft Research
Impact of COVID-19 crisis on the future of work in India
Microsoft Research
Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Microsoft Research
How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
Microsoft Research
Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Microsoft Research
Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Microsoft Research
Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Research
Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Research
Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Research
Directions in ML: Algorithmic foundations of neural architecture search
Microsoft Research
MineRL Competition 2020
Microsoft Research
Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Microsoft Research
From Paper to Product
Microsoft Research
SkinnerDB: Regret Bounded Query Evaluation using RL
Microsoft Research
From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
Microsoft Research
Programming with Proofs for High-assurance Software
Microsoft Research
Platform for Situated Intelligence Overview
Microsoft Research
Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Microsoft Research
Galactic Bell Star Music Demo
Microsoft Research
Importing Animations in Microsoft Expressive Pixels (9 of 9)
Microsoft Research
Welcome to Microsoft Expressive Pixels (1 of 9)
Microsoft Research
Getting Started with Microsoft Expressive Pixels (2 of 9)
Microsoft Research
Creating an Image in Microsoft Expressive Pixels (3 of 9)
Microsoft Research
Creating Animations in Microsoft Expressive Pixels (4 of 9)
Microsoft Research
Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Microsoft Research
Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Microsoft Research
Using Layers in Microsoft Expressive Pixels (7 of 9)
Microsoft Research
Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Microsoft Research
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
Microsoft Research
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
Microsoft Research
Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Microsoft Research
Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Microsoft Research
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Microsoft Research
Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Microsoft Research
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Microsoft Research
Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Microsoft Research
Novel Image Captioning
Microsoft Research
Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Microsoft Research
Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Microsoft Research
How does holographic storage work?
Microsoft Research
The physics of hologram formation in iron doped lithium niobate
Microsoft Research
Introduction to coax: A Modular RL Package
Microsoft Research
Directions in ML: "Neural architecture search: Coming of age"
Microsoft Research
Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research
Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Microsoft Research
Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Microsoft Research
Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research
Clinical Research with FHIR
Microsoft Research
Soundscape Street Preview
Microsoft Research
Tilt-Responsive Techniques for Digital Drawing Boards
Microsoft Research
SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
Microsoft Research
Haptic PIVOT: On-Demand Handhelds in VR
Microsoft Research
SurfaceFleet Supplemental Video Demonstration (UIST 2020)
Microsoft Research
More on: Modern CV Models
View skill →
🎓
Tutor Explanation
DeepCamp AI