Computer Vision Predictions for 2021
Key Takeaways
The Roboflow developer team discusses groundbreaking computer vision developments from 2020 and predicts advancements in 2021, covering topics such as object detection, state-of-the-art models, and multimodal models, with tools like EfficientDet, YOLOv4, and PyTorch.
Full Transcript
[Music] hey there it's joseph brad jacob from roboflow and we're here to talk about computer vision in 2021 so 2020 was a super exciting year in computer vision a lot of new technologies were released and now we're going to kind of reflect on those and look forward to what might be coming in the next year so joseph why don't you start us off with uh some of the things that you were most excited about in 2020 yeah one of the things i was most excited about was the rapid acceleration of state-of-the-art models right so i mean it was in early march that we had efficient debt come out from the google brain team and that broke state of the art in terms of the object detection task and then just one month later uh yolo v4 came out the ov4 paper and what was notable about that was it was the first paper or the first yolo model in the yellow family models that wasn't uh from joseph redmond um right and that's pretty that's pretty notable um and then right after that yellow v5 was released from a different author uh and that i mean doesn't necessarily outperform yellow v4 on the coco benchmark but it seems to localize better to some tasks and so there was this trend of not just being the best on coco but generalizing to other domains and then yet just this last fall scaled yolo e4 came out and this model again broke the state-of-the-art on the coco benchmark and was built using some of the training framework from yolo v5 but with some of the techniques that with csp techniques that improved on the state of the art even more and that was all in 2020 alone so the rate of acceleration it's not just models are getting better but the rate at which they're getting better uh in 2020 was was pretty wide yeah for sure i mean i think the thing that excites me the most is not just that the models are getting better but they're getting easier to use and that means that we're starting to see like all sorts of new use cases and it's not just like google and tesla that are building self-driving cars nowadays you know basically any developer can pick this stuff up off the shelf and build you know whatever they can imagine so we've seen everything from students building uh sushi identifiers to some like fortune 500 companies solving some really challenging problems and these aren't you know tech companies necessarily like we've seen everything from healthcare companies to steel manufacturers and everything in between it's just so cool to to see this technology that used to be a research thing becoming a production thing that's actually affecting our everyday lives yeah yeah that definitely is really exciting i mean there's there's been the coco data set which everyone's been so focused on and the idea was that would lead to generalizable models but now we're actually seeing it happen um and i'm kind of curious why you guys think that is like what do you think uh what were the prerequisites that have made that happen and why that might continue to happen in in 2021 yeah um interestingly i i think um kind of this democratization play um it's been going on for and getting steam for for quite a while now um i this might be a controversial shake i think google collab has actually really kind of propelled it uh you know it used to be you'd have to like rent these expensive gpu machines or build out a box and like that was just kind of a an exclusive sort of club and now kind of anybody can fire up a notebook on google codelab and go through an online course we've now had millions of people go through these massive online courses um with machine learning and computer vision i think you know that just kind of like sparked the the fire and and it's it's really spreading yeah you don't have to take a financial risk to see if it's gonna work now you can just try it out and and see what happens um yeah that's amazing i think accessibility of compute in collab is big i think also the you know that that the release of yolo v5 is kind of a watershed moment in some ways because it was a for the most part kind of like lone wolf person going at it who created a model that challenges state of the art and also is one that was the first family model in the yolo family models that was native to pytorch and so if one dimension of democratization is compute another dimension is certainly ease of use and making that be simpler and so instead of being in a dark net which you know is native to c and it's a bit more challenging to get up and going and yeah you're not aggressively on your first experience that sort the frameworks i've made to be uh adaptable is also why you're seeing tasks not just on the coco data set but i can now try these models on my domain problem totally actually i think that's a great segue into kind of thinking about 2021 and i think um what you described there is actually going to happen also in the deploy side of things so right now you know the the training side of things and creating these prototypes has been democratized but it's still kind of hard to actually get it get your project out there in the wild and you know have people using it and so i think one thing that we're going to see in 2021 is a trend towards making it easy to take these like weights blobs binary blobs that you get out of your model and be able to actually put those into a mobile application or a web application or a hosted api and you know put your project out there um and i think you know this may be a bit self-serving but we saw this the roboflow team last weekend built a hackathon project that served over a hundred thousand people we built it in a weekend with some of these new computer vision models deployed into the cloud with some of these new tools and uh you know we were able to hit scale and and like it wasn't just computer nerds that were playing with our painting game like it was a you know a wide sort of broad broad uh applicability and it was something that you know you didn't need to know anything about the ai that was judging the paintings to be able to use the thing i think that um you know the the deploy is the last missing piece to to really make computer vision this like ubiquitous thing where um you know it's gonna affect everybody's lives not just the developers that are kind of in the know on this new stuff but what do you think i'm on 2021 you have a yeah i mean you have the accessibility is so exciting the training the the deployment but the modeling i think is really going to take off again the biggest thing we saw was that transformers were starting to eat their way into the computer vision space so you know those those had kind of set the standard in in natural language processing but now they're starting to find that you can actually uh just treat an image just as a as a sequence of pixels you know you don't just like you think about a sequence of words um and you know seeing those kind of mass scale transformers start to eat away um so open i released uh the clip model and the a release dolly which are based on these uh transformer models so i think it's gonna be really exciting to see what happens with uh those in the vision space yeah yeah totally so so what is that i mean what would that mean for the capabilities like does that i guess you mentioned dali what what do we see going forward like is that yeah actually transform things that people will notice or is it going to stay in the research realm for now yeah i mean i guess for for the time being probably in research i'm not sure if it's gonna like get released dolly's getting released to everybody or or people are gonna re-engineer it or what but um that's gonna be exciting i mean i might take the opposite side of that i mean if you look at what happened in nlp uh where the cat was out of the bag with i guess was it burp that was first and then you know gbt2 and then gpt3 i think like now that open ai has kind of shown what's possible and that you can steer some of these generative models to to do useful tasks we've already seen you know some hackers trying to steer gans um with the same sort of sort of techniques uh i just think like that my i know my mind got got rolling with all the the possible things that we could build with these new things and i suspect i'm not the only developer that's like uh itching to build stuff with it so i think even if openad doesn't release their model i would be surprised if by the end of the year we don't have some sort of steerable generative model that can can produce useful stuff yeah do you have a dark horse for for 2021's biggest innovation coming down the pipeline or a theme to watch i think the i mean i was gonna i was gonna take the side that you were with the transformers eating vision and almost having like a image to back understanding instead of having to like re-embed and uniquely understand even each respective domain problem i mean one thing that i'd like to see i think that could be interesting to see is the creation of a new coco benchmark in some senses or a new way to measure generalizability of models or um so i mean you mentioned clip that was trained in 400 million image text pairs right and the cocoa data said it's only in the hundreds of thousands low hundreds of thousands and i think that um some of the limitations we saw in 2020 was almost debates about what is state of the art because it becomes domain dependent and so i think what would be really powerful is to see um maybe like a a benchmark that measures domain adaptation uh of a model um or maybe that's not requisite if you have like this image to vec idea where you don't have to retrain it for each domain but something that i'm excited to work on with roboflow and the platform that we have to make that happen is public.roboflow.com being a place where people can start to host their data sets per domain and so you can almost like take you know the off-the-shelf like tennis ball model and adapt that better to like what your scene is of tennis balls that's sort of like a micro like empowering the every person sort of perspective but i could also see you know microsoft would release the original cocoa data set in 2014 coming out with another someone like that coming out with another massive uh data set for benchmarking or even industry benchmarking um i don't know that i'd place money on this happening but i think we're overdue for it at a minimum right like we've been bench marketing it's the same thing for for so long so yeah that's one thing i'd hoped for i think it's interesting talking about cocoa specifically um is if you you know kind of dive into some of those images like there's a lot of errors in that data set and so we've been kind of like you know benchmarking off of this not perfect data set and we're you know are we going to get to the point where the models are so good that they're better than the human amputators were when that data set came out you know almost a decade ago um and i think we've see we've seen that with mnist too if you look at you know those models that are getting 98 99 accuracy on mnist you look at the ones that they missed and actually i think the models probably are more accurate than the humans you labeled on right and so you you kind of tap out you know what what that even tells you um if the model is doing better than the actual accuracy so i think that's right and i think generalizability is is a huge thing you know it's one thing to do well on coco it's another thing to do well on people's actual real world problems um so yeah i i think i'd agree with that i would uh be interested in more discussion around what what is it that you know a set of ideal data sets would have to look like to to make them be the new you know standard that researchers are pinning against to make their research be applicable to actual real world situations what problems do you think are not going to get solved in 2021 we're talking about a lot of things that we're really like excited about and domain adaptation and new benchmarks and transformers and reducing training costs and accessibility but like i i mean like i'm not i suspect like some of the low-hanging fruit there there's a you know a high neighborhood of possibilities or things that aren't gonna we're not gonna see a general ai maybe you're maybe you think we're ready for it in the next year but i would take that we're probably not ready for general ai but in terms of things that like might be disappointing or places that like we're going to see slower progress what are going to be problems that are still outstanding if i had to pose this question what are areas where there's going to be more work to do well so on the deploy side i think that's always going to be somewhat of a frustration of uh having to get all the dependencies together to stand up a model and get it to run of course you have exciting things starting to happen with like onyx kind of trying to make a unified model framework but i think this this problem of getting a model that you think is working and then trying to get it actually into the production use case running at the fps that you wanted to um we'll probably persist progress on it for sure like that that will be with us for 2021 but maybe not too much longer yeah i i think it's gonna be a while before we get um any sort of you know really generalized computer vision model you know not not even talking about a generalized ai but um you know we're starting to see beginnings of like what it would look like for an you know unsupervised computer vision world but i think it's going to take actually quite a bit of time for that to come and subsume custom models and so i think we'll continue to see you know there being an importance of you know collecting a representative data set and getting it labeled by human annotators and trading a custom model with custom weights i think you know if you look at decade down the line perhaps um you know the models get good enough that that isn't needed but i think you know if we're talking about what what is you know not going to change in 2021 i would be very surprised if you know we're sitting here a year from now and we're saying hey we have this you know supermodel that that does everything right one thing that i'm thinking about too is um everyone that subscribes to this channel or if you're watching this you should be a subscriber to this channel that might miss out on something is um reaching the uh one level abstracted audience of users of computer vision so like i think that like in 2020 we saw a lot of momentum a lot of deep momentum and and researchers and those that are really domain experts and those researchers are going to be are already firing in all cylinders in 2021 i mean open ai's models technically came out january 5 of this year um in some senses in 2021 so i think there but there's also these other audiences of like if i'm a developer like how much about convolution do i need to know or even if i'm like one level higher than that if i'm a um i don't mean higher in terms of like hierarchy i mean in terms of level of abstraction like if i am a technical product manager and i mean i'm not writing code maybe all the time but i certainly have a grasp of these things or even maybe like someone that has very little knowledge of ml i think that like 2021 we're gonna start to tap into some of those personas other personas but i still like it's remarkable how much um in some senses like expertise on the fringes visionist generally and i don't see that being solved in 2021 i see this starting to tap into that but there's so many uh uh subject matter experts that can benefit from this technology that like those that are deep in the field haven't even considered how to uh help yeah i think there's a big challenge there too with you know educating you want those domain experts to know enough to build um really cool things but the the challenge is you also need to give them the tools and the knowledge that they need to not fall into common traps and so i think like you know as we see continue to see um this sort of um expansion of machine learning from the data science realm to developers into to domain experts things like you know bias in your models are things that are going to become something that needs to be talked about more generally and i think that you know in 2021 especially as we you know start to see a lot of these people using these models for the first time um i think that a lot of people are going to trip over themselves with some of those sorts of things and so it'll be interesting to see what we can do to educate people um and maybe it's just going to have to be something where you know there there's a pr thing around it and it gets more widespread um understanding beyond just the you know the data scientists that have been thinking about that sort of stuff for a long time um you know we saw that with i think it was just a couple years ago the apple card and um you know the the credit score between men and women and and the credit limits but i think that a lot of those same sorts of things are going to apply in the vision realm and we've seen a bit of it with like facial recognition stuff already but i think there's probably a lot of corner cases that folks haven't even discovered or thought about yet that as vision expands into every single domain are going to have to be addressed and be a part of that basic you know education of coming into this data science world for the first time as a non-data scientist like what are the things that i need to know and watch out for and then hopefully tools that will automatically help people not fall into those same traps makes a lot of sense definitely definitely cool do you have a uh let's let's uh let's say a less than five percent probability and greater than one percent like what what's the outside chance of something that might happen in 2021 but it's unlikely that that you would uh or that you think might be underrated in terms of likelihood um well we talked about image to back but i think that's gonna be more likely uh than five percent um i i would probably go with um some sort of uh zero shot uh object detection model somewhere in the one to five yeah yeah i guess i painted myself in a corner here because i again i think this is probably gonna actually be more than five percent but along those lines like multimodal models where they they take knowledge that is not just what's in the pixel data yeah um and use you know outside knowledge either from what they've learned from language um or potentially you know other things that they've gathered about metadata or those sorts of things um i i think you know we'll start to see some like if there was a yellow v5 plus metadata model or a yellow v5 plus text model that came out i wouldn't be surprised i'd be slightly surprised to see it this year but only because it would be like a year early um but i think that would be cool to see so people are working so you're thinking you you input the image along with some text features right yeah and that helps the model you know understand versus just sending in a you know a bunch of pixels with no context yeah certainly i think it's similar to those lines but something that's been remarkably underserved in my head is video just being treated as images and i think that now we have sufficient compute now there's enough video there's no shortage of data and the modeling techniques are getting there so this isn't probably in the family of multi-modal models because i would imagine that the way that you solve modeling video differently would be with something of the sequencing of frames something like an rnn for video yeah and rnn for video right because like if you have a model that's doing video inference and it misses something in one of the frames now right now we're really naive and we're like oh wow like you know frames one to seven had this thing eight didn't but nine and ten did eight probably also had that and the frame information from the surrounding frame should inform that and so i think it's i don't know maybe a one to five percent chance that i think we have the state-of-the-art video architecture model comes out um and that's unique and different um i think it's going to happen inevitably but maybe not 2021 that has callbacks to what you were talking about earlier with the transformer models coming into the the vision realm maybe as jacob that was saying that i think you know those those transformer models seem to do quite well with you know data sequences yeah where you could add you know the temporal dimension and i can see that being um groundbreaking and getting to something that is less of just a sequence of frames and more of a cohesive video cool well we we're super excited to see what comes out in 2021 and uh if you have predictions or ideas for what you think is going to come out that we didn't didn't come up with or you strongly disagree with one of our predictions uh we would love to hear it in the comments below uh we'll be sure to fan the flames of that of that play more down there and yeah until next time uh this has been a roboflow fireside chat uh we'll uh we'll talk to you later be sure to like and subscribe
Original Description
The Roboflow developer team recaps the groundbreaking computer vision developments from 2020 and speculates about what we might expect from the research and development community in 2021.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Roboflow · Roboflow · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
YOLOv3 PyTorch Notebook Tutorial
Roboflow
How to Train YOLOv4 on a Custom Dataset (PyTorch)
Roboflow
How to Train YOLOv5 on a Custom Dataset
Roboflow
How to Use the Roboflow Dataset Health Check
Roboflow
What is Mean Average Precision (mAP)?
Roboflow
How to Use the Roboflow Model Library
Roboflow
How to Train EfficientDet in TensorFlow 2 Object Detection
Roboflow
How to Train YOLO v4 Tiny (Darknet) on a Custom Dataset
Roboflow
Ask the Roboflow Team Anything - Episode 1
Roboflow
Exploring The COCO Dataset
Roboflow
Community Spotlight: Improving Uno with Computer Vision
Roboflow
Mosaic Data Augmentation - Deep Dive
Roboflow
Hands on with the OAK-1
Roboflow
Glenn Jocher: What is New in YOLO v5?
Roboflow
How to Use Amazon Rekognition Custom Labels and Roboflow to Build an Object Detection Model
Roboflow
An Interview with Brandon Gilles, Luxonis Founder and OAK Chief Architect
Roboflow
How to Train a Custom Mobile Object Detection Model (with YOLOv4 Tiny and TensorFlow Lite)
Roboflow
Tackling the Small Object Problem in Object Detection
Roboflow
Fast.ai v2 Released - What's New?
Roboflow
Teaser: Roboflow Train (1-Click Computer Vision AutoML)
Roboflow
How to Train a Custom Resnet34 Image Classification Model
Roboflow
How to Label Images for Object Detection with CVAT
Roboflow
Deploy YOLOv5 to Jetson Xavier NX at 30 FPS
Roboflow
Elisha Odemakinde Hosts Roboflow ML Engineer, Jacob Solawetz
Roboflow
Getting Started with VoTT - Computer Vision Annotation
Roboflow
How to Manage Classes in Object Detection (Rename, Combine, Balance)
Roboflow
How to Train YOLOv4 on a Custom Dataset in Darknet
Roboflow
Is Grayscale a Preprocessing or Augmentation Step in Computer Vision?
Roboflow
Getting Started with Image Data Augmentation
Roboflow
Glenn Jocher: Image Augmentation in YOLO v5 and Beyond
Roboflow
GA Hosts Roboflow - Healthcare and AI
Roboflow
How do self driving cars know when to stop?
Roboflow
What is PASCAL VOC XML?
Roboflow
AutoML Showdown: Google vs Amazon vs Microsoft
Roboflow
How is computer vision changing manufacturing?
Roboflow
The Alphabet in American Sign Language
Roboflow
Luxonis OAK-D: Computer Vision on Device
Roboflow
How to Train a Custom Faster R-CNN Model with Facebook AI's Detectron2 | Use Your Own Dataset
Roboflow
TensorFlow vs PyTorch: Fireside
Roboflow
Occlusion Techniques in Computer Vision
Roboflow
A Customizable Web Application for Your Computer Vision Model
Roboflow
Model Tradeoffs and the Future of Computer Vision
Roboflow
Designing an Augmented Reality Board Game App
Roboflow
YOLOv4 - Advanced Tactics
Roboflow
How to Use CreateML and Build a Computer Vision iPhone App | AR Object Detection
Roboflow
Fireside Chat: Computer Vision in Agriculture
Roboflow
Scaled-YOLOv4 Tops EfficientDet: Research Rundown
Roboflow
What is Image Preprocessing?
Roboflow
Building a Community of Creators with BlkArthouse and Von Deon
Roboflow
How to Train Scaled-YOLOv4 to Detect Custom Objects
Roboflow
Intro to Computer Vision: Fireside
Roboflow
The Best Way to Annotate Images for Object Detection
Roboflow
The Computer Vision Process: Fireside
Roboflow
How to Annotate Images with Your Team Using Roboflow
Roboflow
Introducing the Roboflow Object Count Histogram
Roboflow
How Fast is the M1 at Machine Learning? Benchmarking Apple's M1 and Intel's Chips
Roboflow
CLIP: OpenAI's amazing new zero-shot image classifier
Roboflow
How I hacked my Nest camera to run custom models
Roboflow
Getting Started with the Roboflow Inference API
Roboflow
Transfer Learning in Computer Vision | What, How, Why
Roboflow
More on: Modern CV Models
View skill →
🎓
Tutor Explanation
DeepCamp AI