Computer Vision Predictions for 2021

Roboflow · Advanced ·👁️ Computer Vision ·5y ago

Key Takeaways

The Roboflow developer team discusses groundbreaking computer vision developments from 2020 and predicts advancements in 2021, covering topics such as object detection, state-of-the-art models, and multimodal models, with tools like EfficientDet, YOLOv4, and PyTorch.

Full Transcript

[Music] hey there it's joseph brad jacob from roboflow and we're here to talk about computer vision in 2021 so 2020 was a super exciting year in computer vision a lot of new technologies were released and now we're going to kind of reflect on those and look forward to what might be coming in the next year so joseph why don't you start us off with uh some of the things that you were most excited about in 2020 yeah one of the things i was most excited about was the rapid acceleration of state-of-the-art models right so i mean it was in early march that we had efficient debt come out from the google brain team and that broke state of the art in terms of the object detection task and then just one month later uh yolo v4 came out the ov4 paper and what was notable about that was it was the first paper or the first yolo model in the yellow family models that wasn't uh from joseph redmond um right and that's pretty that's pretty notable um and then right after that yellow v5 was released from a different author uh and that i mean doesn't necessarily outperform yellow v4 on the coco benchmark but it seems to localize better to some tasks and so there was this trend of not just being the best on coco but generalizing to other domains and then yet just this last fall scaled yolo e4 came out and this model again broke the state-of-the-art on the coco benchmark and was built using some of the training framework from yolo v5 but with some of the techniques that with csp techniques that improved on the state of the art even more and that was all in 2020 alone so the rate of acceleration it's not just models are getting better but the rate at which they're getting better uh in 2020 was was pretty wide yeah for sure i mean i think the thing that excites me the most is not just that the models are getting better but they're getting easier to use and that means that we're starting to see like all sorts of new use cases and it's not just like google and tesla that are building self-driving cars nowadays you know basically any developer can pick this stuff up off the shelf and build you know whatever they can imagine so we've seen everything from students building uh sushi identifiers to some like fortune 500 companies solving some really challenging problems and these aren't you know tech companies necessarily like we've seen everything from healthcare companies to steel manufacturers and everything in between it's just so cool to to see this technology that used to be a research thing becoming a production thing that's actually affecting our everyday lives yeah yeah that definitely is really exciting i mean there's there's been the coco data set which everyone's been so focused on and the idea was that would lead to generalizable models but now we're actually seeing it happen um and i'm kind of curious why you guys think that is like what do you think uh what were the prerequisites that have made that happen and why that might continue to happen in in 2021 yeah um interestingly i i think um kind of this democratization play um it's been going on for and getting steam for for quite a while now um i this might be a controversial shake i think google collab has actually really kind of propelled it uh you know it used to be you'd have to like rent these expensive gpu machines or build out a box and like that was just kind of a an exclusive sort of club and now kind of anybody can fire up a notebook on google codelab and go through an online course we've now had millions of people go through these massive online courses um with machine learning and computer vision i think you know that just kind of like sparked the the fire and and it's it's really spreading yeah you don't have to take a financial risk to see if it's gonna work now you can just try it out and and see what happens um yeah that's amazing i think accessibility of compute in collab is big i think also the you know that that the release of yolo v5 is kind of a watershed moment in some ways because it was a for the most part kind of like lone wolf person going at it who created a model that challenges state of the art and also is one that was the first family model in the yolo family models that was native to pytorch and so if one dimension of democratization is compute another dimension is certainly ease of use and making that be simpler and so instead of being in a dark net which you know is native to c and it's a bit more challenging to get up and going and yeah you're not aggressively on your first experience that sort the frameworks i've made to be uh adaptable is also why you're seeing tasks not just on the coco data set but i can now try these models on my domain problem totally actually i think that's a great segue into kind of thinking about 2021 and i think um what you described there is actually going to happen also in the deploy side of things so right now you know the the training side of things and creating these prototypes has been democratized but it's still kind of hard to actually get it get your project out there in the wild and you know have people using it and so i think one thing that we're going to see in 2021 is a trend towards making it easy to take these like weights blobs binary blobs that you get out of your model and be able to actually put those into a mobile application or a web application or a hosted api and you know put your project out there um and i think you know this may be a bit self-serving but we saw this the roboflow team last weekend built a hackathon project that served over a hundred thousand people we built it in a weekend with some of these new computer vision models deployed into the cloud with some of these new tools and uh you know we were able to hit scale and and like it wasn't just computer nerds that were playing with our painting game like it was a you know a wide sort of broad broad uh applicability and it was something that you know you didn't need to know anything about the ai that was judging the paintings to be able to use the thing i think that um you know the the deploy is the last missing piece to to really make computer vision this like ubiquitous thing where um you know it's gonna affect everybody's lives not just the developers that are kind of in the know on this new stuff but what do you think i'm on 2021 you have a yeah i mean you have the accessibility is so exciting the training the the deployment but the modeling i think is really going to take off again the biggest thing we saw was that transformers were starting to eat their way into the computer vision space so you know those those had kind of set the standard in in natural language processing but now they're starting to find that you can actually uh just treat an image just as a as a sequence of pixels you know you don't just like you think about a sequence of words um and you know seeing those kind of mass scale transformers start to eat away um so open i released uh the clip model and the a release dolly which are based on these uh transformer models so i think it's gonna be really exciting to see what happens with uh those in the vision space yeah yeah totally so so what is that i mean what would that mean for the capabilities like does that i guess you mentioned dali what what do we see going forward like is that yeah actually transform things that people will notice or is it going to stay in the research realm for now yeah i mean i guess for for the time being probably in research i'm not sure if it's gonna like get released dolly's getting released to everybody or or people are gonna re-engineer it or what but um that's gonna be exciting i mean i might take the opposite side of that i mean if you look at what happened in nlp uh where the cat was out of the bag with i guess was it burp that was first and then you know gbt2 and then gpt3 i think like now that open ai has kind of shown what's possible and that you can steer some of these generative models to to do useful tasks we've already seen you know some hackers trying to steer gans um with the same sort of sort of techniques uh i just think like that my i know my mind got got rolling with all the the possible things that we could build with these new things and i suspect i'm not the only developer that's like uh itching to build stuff with it so i think even if openad doesn't release their model i would be surprised if by the end of the year we don't have some sort of steerable generative model that can can produce useful stuff yeah do you have a dark horse for for 2021's biggest innovation coming down the pipeline or a theme to watch i think the i mean i was gonna i was gonna take the side that you were with the transformers eating vision and almost having like a image to back understanding instead of having to like re-embed and uniquely understand even each respective domain problem i mean one thing that i'd like to see i think that could be interesting to see is the creation of a new coco benchmark in some senses or a new way to measure generalizability of models or um so i mean you mentioned clip that was trained in 400 million image text pairs right and the cocoa data said it's only in the hundreds of thousands low hundreds of thousands and i think that um some of the limitations we saw in 2020 was almost debates about what is state of the art because it becomes domain dependent and so i think what would be really powerful is to see um maybe like a a benchmark that measures domain adaptation uh of a model um or maybe that's not requisite if you have like this image to vec idea where you don't have to retrain it for each domain but something that i'm excited to work on with roboflow and the platform that we have to make that happen is public.roboflow.com being a place where people can start to host their data sets per domain and so you can almost like take you know the off-the-shelf like tennis ball model and adapt that better to like what your scene is of tennis balls that's sort of like a micro like empowering the every person sort of perspective but i could also see you know microsoft would release the original cocoa data set in 2014 coming out with another someone like that coming out with another massive uh data set for benchmarking or even industry benchmarking um i don't know that i'd place money on this happening but i think we're overdue for it at a minimum right like we've been bench marketing it's the same thing for for so long so yeah that's one thing i'd hoped for i think it's interesting talking about cocoa specifically um is if you you know kind of dive into some of those images like there's a lot of errors in that data set and so we've been kind of like you know benchmarking off of this not perfect data set and we're you know are we going to get to the point where the models are so good that they're better than the human amputators were when that data set came out you know almost a decade ago um and i think we've see we've seen that with mnist too if you look at you know those models that are getting 98 99 accuracy on mnist you look at the ones that they missed and actually i think the models probably are more accurate than the humans you labeled on right and so you you kind of tap out you know what what that even tells you um if the model is doing better than the actual accuracy so i think that's right and i think generalizability is is a huge thing you know it's one thing to do well on coco it's another thing to do well on people's actual real world problems um so yeah i i think i'd agree with that i would uh be interested in more discussion around what what is it that you know a set of ideal data sets would have to look like to to make them be the new you know standard that researchers are pinning against to make their research be applicable to actual real world situations what problems do you think are not going to get solved in 2021 we're talking about a lot of things that we're really like excited about and domain adaptation and new benchmarks and transformers and reducing training costs and accessibility but like i i mean like i'm not i suspect like some of the low-hanging fruit there there's a you know a high neighborhood of possibilities or things that aren't gonna we're not gonna see a general ai maybe you're maybe you think we're ready for it in the next year but i would take that we're probably not ready for general ai but in terms of things that like might be disappointing or places that like we're going to see slower progress what are going to be problems that are still outstanding if i had to pose this question what are areas where there's going to be more work to do well so on the deploy side i think that's always going to be somewhat of a frustration of uh having to get all the dependencies together to stand up a model and get it to run of course you have exciting things starting to happen with like onyx kind of trying to make a unified model framework but i think this this problem of getting a model that you think is working and then trying to get it actually into the production use case running at the fps that you wanted to um we'll probably persist progress on it for sure like that that will be with us for 2021 but maybe not too much longer yeah i i think it's gonna be a while before we get um any sort of you know really generalized computer vision model you know not not even talking about a generalized ai but um you know we're starting to see beginnings of like what it would look like for an you know unsupervised computer vision world but i think it's going to take actually quite a bit of time for that to come and subsume custom models and so i think we'll continue to see you know there being an importance of you know collecting a representative data set and getting it labeled by human annotators and trading a custom model with custom weights i think you know if you look at decade down the line perhaps um you know the models get good enough that that isn't needed but i think you know if we're talking about what what is you know not going to change in 2021 i would be very surprised if you know we're sitting here a year from now and we're saying hey we have this you know supermodel that that does everything right one thing that i'm thinking about too is um everyone that subscribes to this channel or if you're watching this you should be a subscriber to this channel that might miss out on something is um reaching the uh one level abstracted audience of users of computer vision so like i think that like in 2020 we saw a lot of momentum a lot of deep momentum and and researchers and those that are really domain experts and those researchers are going to be are already firing in all cylinders in 2021 i mean open ai's models technically came out january 5 of this year um in some senses in 2021 so i think there but there's also these other audiences of like if i'm a developer like how much about convolution do i need to know or even if i'm like one level higher than that if i'm a um i don't mean higher in terms of like hierarchy i mean in terms of level of abstraction like if i am a technical product manager and i mean i'm not writing code maybe all the time but i certainly have a grasp of these things or even maybe like someone that has very little knowledge of ml i think that like 2021 we're gonna start to tap into some of those personas other personas but i still like it's remarkable how much um in some senses like expertise on the fringes visionist generally and i don't see that being solved in 2021 i see this starting to tap into that but there's so many uh uh subject matter experts that can benefit from this technology that like those that are deep in the field haven't even considered how to uh help yeah i think there's a big challenge there too with you know educating you want those domain experts to know enough to build um really cool things but the the challenge is you also need to give them the tools and the knowledge that they need to not fall into common traps and so i think like you know as we see continue to see um this sort of um expansion of machine learning from the data science realm to developers into to domain experts things like you know bias in your models are things that are going to become something that needs to be talked about more generally and i think that you know in 2021 especially as we you know start to see a lot of these people using these models for the first time um i think that a lot of people are going to trip over themselves with some of those sorts of things and so it'll be interesting to see what we can do to educate people um and maybe it's just going to have to be something where you know there there's a pr thing around it and it gets more widespread um understanding beyond just the you know the data scientists that have been thinking about that sort of stuff for a long time um you know we saw that with i think it was just a couple years ago the apple card and um you know the the credit score between men and women and and the credit limits but i think that a lot of those same sorts of things are going to apply in the vision realm and we've seen a bit of it with like facial recognition stuff already but i think there's probably a lot of corner cases that folks haven't even discovered or thought about yet that as vision expands into every single domain are going to have to be addressed and be a part of that basic you know education of coming into this data science world for the first time as a non-data scientist like what are the things that i need to know and watch out for and then hopefully tools that will automatically help people not fall into those same traps makes a lot of sense definitely definitely cool do you have a uh let's let's uh let's say a less than five percent probability and greater than one percent like what what's the outside chance of something that might happen in 2021 but it's unlikely that that you would uh or that you think might be underrated in terms of likelihood um well we talked about image to back but i think that's gonna be more likely uh than five percent um i i would probably go with um some sort of uh zero shot uh object detection model somewhere in the one to five yeah yeah i guess i painted myself in a corner here because i again i think this is probably gonna actually be more than five percent but along those lines like multimodal models where they they take knowledge that is not just what's in the pixel data yeah um and use you know outside knowledge either from what they've learned from language um or potentially you know other things that they've gathered about metadata or those sorts of things um i i think you know we'll start to see some like if there was a yellow v5 plus metadata model or a yellow v5 plus text model that came out i wouldn't be surprised i'd be slightly surprised to see it this year but only because it would be like a year early um but i think that would be cool to see so people are working so you're thinking you you input the image along with some text features right yeah and that helps the model you know understand versus just sending in a you know a bunch of pixels with no context yeah certainly i think it's similar to those lines but something that's been remarkably underserved in my head is video just being treated as images and i think that now we have sufficient compute now there's enough video there's no shortage of data and the modeling techniques are getting there so this isn't probably in the family of multi-modal models because i would imagine that the way that you solve modeling video differently would be with something of the sequencing of frames something like an rnn for video yeah and rnn for video right because like if you have a model that's doing video inference and it misses something in one of the frames now right now we're really naive and we're like oh wow like you know frames one to seven had this thing eight didn't but nine and ten did eight probably also had that and the frame information from the surrounding frame should inform that and so i think it's i don't know maybe a one to five percent chance that i think we have the state-of-the-art video architecture model comes out um and that's unique and different um i think it's going to happen inevitably but maybe not 2021 that has callbacks to what you were talking about earlier with the transformer models coming into the the vision realm maybe as jacob that was saying that i think you know those those transformer models seem to do quite well with you know data sequences yeah where you could add you know the temporal dimension and i can see that being um groundbreaking and getting to something that is less of just a sequence of frames and more of a cohesive video cool well we we're super excited to see what comes out in 2021 and uh if you have predictions or ideas for what you think is going to come out that we didn't didn't come up with or you strongly disagree with one of our predictions uh we would love to hear it in the comments below uh we'll be sure to fan the flames of that of that play more down there and yeah until next time uh this has been a roboflow fireside chat uh we'll uh we'll talk to you later be sure to like and subscribe

Original Description

The Roboflow developer team recaps the groundbreaking computer vision developments from 2020 and speculates about what we might expect from the research and development community in 2021.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Roboflow · Roboflow · 0 of 60

← Previous Next →
1 YOLOv3 PyTorch Notebook Tutorial
YOLOv3 PyTorch Notebook Tutorial
Roboflow
2 How to Train YOLOv4 on a Custom Dataset (PyTorch)
How to Train YOLOv4 on a Custom Dataset (PyTorch)
Roboflow
3 How to Train YOLOv5 on a Custom Dataset
How to Train YOLOv5 on a Custom Dataset
Roboflow
4 How to Use the Roboflow Dataset Health Check
How to Use the Roboflow Dataset Health Check
Roboflow
5 What is Mean Average Precision (mAP)?
What is Mean Average Precision (mAP)?
Roboflow
6 How to Use the Roboflow Model Library
How to Use the Roboflow Model Library
Roboflow
7 How to Train EfficientDet in TensorFlow 2 Object Detection
How to Train EfficientDet in TensorFlow 2 Object Detection
Roboflow
8 How to Train YOLO v4 Tiny (Darknet) on a Custom Dataset
How to Train YOLO v4 Tiny (Darknet) on a Custom Dataset
Roboflow
9 Ask the Roboflow Team Anything - Episode 1
Ask the Roboflow Team Anything - Episode 1
Roboflow
10 Exploring The COCO Dataset
Exploring The COCO Dataset
Roboflow
11 Community Spotlight: Improving Uno with Computer Vision
Community Spotlight: Improving Uno with Computer Vision
Roboflow
12 Mosaic Data Augmentation - Deep Dive
Mosaic Data Augmentation - Deep Dive
Roboflow
13 Hands on with the OAK-1
Hands on with the OAK-1
Roboflow
14 Glenn Jocher: What is New in YOLO v5?
Glenn Jocher: What is New in YOLO v5?
Roboflow
15 How to Use Amazon Rekognition Custom Labels and Roboflow to Build an Object Detection Model
How to Use Amazon Rekognition Custom Labels and Roboflow to Build an Object Detection Model
Roboflow
16 An Interview with Brandon Gilles, Luxonis Founder and OAK Chief Architect
An Interview with Brandon Gilles, Luxonis Founder and OAK Chief Architect
Roboflow
17 How to Train a Custom Mobile Object Detection Model (with YOLOv4 Tiny and TensorFlow Lite)
How to Train a Custom Mobile Object Detection Model (with YOLOv4 Tiny and TensorFlow Lite)
Roboflow
18 Tackling the Small Object Problem in Object Detection
Tackling the Small Object Problem in Object Detection
Roboflow
19 Fast.ai v2 Released - What's New?
Fast.ai v2 Released - What's New?
Roboflow
20 Teaser: Roboflow Train (1-Click Computer Vision AutoML)
Teaser: Roboflow Train (1-Click Computer Vision AutoML)
Roboflow
21 How to Train a Custom Resnet34 Image Classification Model
How to Train a Custom Resnet34 Image Classification Model
Roboflow
22 How to Label Images for Object Detection with CVAT
How to Label Images for Object Detection with CVAT
Roboflow
23 Deploy YOLOv5 to Jetson Xavier NX at 30 FPS
Deploy YOLOv5 to Jetson Xavier NX at 30 FPS
Roboflow
24 Elisha Odemakinde Hosts Roboflow ML Engineer, Jacob Solawetz
Elisha Odemakinde Hosts Roboflow ML Engineer, Jacob Solawetz
Roboflow
25 Getting Started with VoTT - Computer Vision Annotation
Getting Started with VoTT - Computer Vision Annotation
Roboflow
26 How to Manage Classes in Object Detection (Rename, Combine, Balance)
How to Manage Classes in Object Detection (Rename, Combine, Balance)
Roboflow
27 How to Train YOLOv4 on a Custom Dataset in Darknet
How to Train YOLOv4 on a Custom Dataset in Darknet
Roboflow
28 Is Grayscale a Preprocessing or Augmentation Step in Computer Vision?
Is Grayscale a Preprocessing or Augmentation Step in Computer Vision?
Roboflow
29 Getting Started with Image Data Augmentation
Getting Started with Image Data Augmentation
Roboflow
30 Glenn Jocher: Image Augmentation in YOLO v5 and Beyond
Glenn Jocher: Image Augmentation in YOLO v5 and Beyond
Roboflow
31 GA Hosts Roboflow - Healthcare and AI
GA Hosts Roboflow - Healthcare and AI
Roboflow
32 How do self driving cars know when to stop?
How do self driving cars know when to stop?
Roboflow
33 What is PASCAL VOC XML?
What is PASCAL VOC XML?
Roboflow
34 AutoML Showdown: Google vs Amazon vs Microsoft
AutoML Showdown: Google vs Amazon vs Microsoft
Roboflow
35 How is computer vision changing manufacturing?
How is computer vision changing manufacturing?
Roboflow
36 The Alphabet in American Sign Language
The Alphabet in American Sign Language
Roboflow
37 Luxonis OAK-D: Computer Vision on Device
Luxonis OAK-D: Computer Vision on Device
Roboflow
38 How to Train a Custom Faster R-CNN Model with Facebook AI's Detectron2 | Use Your Own Dataset
How to Train a Custom Faster R-CNN Model with Facebook AI's Detectron2 | Use Your Own Dataset
Roboflow
39 TensorFlow vs PyTorch: Fireside
TensorFlow vs PyTorch: Fireside
Roboflow
40 Occlusion Techniques in Computer Vision
Occlusion Techniques in Computer Vision
Roboflow
41 A Customizable Web Application for Your Computer Vision Model
A Customizable Web Application for Your Computer Vision Model
Roboflow
42 Model Tradeoffs and the Future of Computer Vision
Model Tradeoffs and the Future of Computer Vision
Roboflow
43 Designing an Augmented Reality Board Game App
Designing an Augmented Reality Board Game App
Roboflow
44 YOLOv4 - Advanced Tactics
YOLOv4 - Advanced Tactics
Roboflow
45 How to Use CreateML and Build a Computer Vision iPhone App | AR Object Detection
How to Use CreateML and Build a Computer Vision iPhone App | AR Object Detection
Roboflow
46 Fireside Chat: Computer Vision in Agriculture
Fireside Chat: Computer Vision in Agriculture
Roboflow
47 Scaled-YOLOv4 Tops EfficientDet: Research Rundown
Scaled-YOLOv4 Tops EfficientDet: Research Rundown
Roboflow
48 What is Image Preprocessing?
What is Image Preprocessing?
Roboflow
49 Building a Community of Creators with BlkArthouse and Von Deon
Building a Community of Creators with BlkArthouse and Von Deon
Roboflow
50 How to Train Scaled-YOLOv4 to Detect Custom Objects
How to Train Scaled-YOLOv4 to Detect Custom Objects
Roboflow
51 Intro to Computer Vision: Fireside
Intro to Computer Vision: Fireside
Roboflow
52 The Best Way to Annotate Images for Object Detection
The Best Way to Annotate Images for Object Detection
Roboflow
53 The Computer Vision Process: Fireside
The Computer Vision Process: Fireside
Roboflow
54 How to Annotate Images with Your Team Using Roboflow
How to Annotate Images with Your Team Using Roboflow
Roboflow
55 Introducing the Roboflow Object Count Histogram
Introducing the Roboflow Object Count Histogram
Roboflow
56 How Fast is the M1 at Machine Learning? Benchmarking Apple's M1 and Intel's Chips
How Fast is the M1 at Machine Learning? Benchmarking Apple's M1 and Intel's Chips
Roboflow
57 CLIP: OpenAI's amazing new zero-shot image classifier
CLIP: OpenAI's amazing new zero-shot image classifier
Roboflow
58 How I hacked my Nest camera to run custom models
How I hacked my Nest camera to run custom models
Roboflow
59 Getting Started with the Roboflow Inference API
Getting Started with the Roboflow Inference API
Roboflow
60 Transfer Learning in Computer Vision | What, How, Why
Transfer Learning in Computer Vision | What, How, Why
Roboflow

The video discusses the latest advancements in computer vision, including object detection, state-of-the-art models, and multimodal models, and predicts future developments in 2021. Viewers can learn about the latest tools and techniques in computer vision and how to apply them in practice.

Key Takeaways
  1. Explore EfficientDet and YOLOv4 for object detection
  2. Use PyTorch for computer vision tasks
  3. Investigate DALL-E for generative tasks
  4. Learn about multimodal models and their applications
  5. Understand the importance of domain adaptation and benchmarking
💡 The field of computer vision is rapidly advancing, with new state-of-the-art models and techniques being developed, and multimodal models are becoming increasingly important.

Related Reads

Up next
9-Phase Computer Vision Roadmap 2026 | AI & Deep Learning | #shorts
SCALER
Watch →