Focal Transformer: Focal Self-attention for Local-Global Interactions in Vision Transformers

Aleksa Gordić - The AI Epiphany · Beginner ·🧠 Large Language Models ·5y ago

Key Takeaways

This video covers the Focal Transformer paper, introducing focal self-attention for local-global interactions in vision transformers

Full Transcript

What's up? In this video I'm covering focal self-attention for local global interactions in vision transformers by the Microsoft research and Microsoft cloud and AI teams. So they basically introduced this novel layer called focal self-attention and introduced this novel model called focal transformer. And let's see what it's all about. Basically, if you focus on this diagram here, let me kind of motivate the introduction of this focal self-attention. First things first, on the left-hand side we have a classical vision transformer, so that's date model from Facebook AI. And if you haven't watched my video on vision transformer, I do suggest you go ahead and watch it first, but I'll just briefly explain how it works here. So the idea is because you have a transformer and transformers were invented in the context of NLP where you introduce like a sequence of tokens as the input, so we need we somehow need to convert this image into that sequence of tokens. And so we So how we do that is basically we split the image into patches, so something like this. So we'll have a patch here, we have a patch here, etc. So the whole image will be basically built from these patches. And here you can see a query patch. So this is one particular patch and obviously hopefully you're you're familiar with the self-attention like like terminology, but this is a query patch and what we can see on the on this these attention maps here is basically that the vision transformer has both local attention as well as global attention. So let me like kind of explain what I mean by that. So if you focus on this on this attention map here. So what it basically says is the following. That means that as you can see this like blob here, that means that this query patch is basically attending these nearby patches a lot. And that's just one particular attention head from the first layer of this of this state transformer. And if we focus on the second head or third head, we can see that it also has So this query patch in those other heads also tends certain other areas. Like if you focus on this one, you can see that that part is these patches here. And it's also focusing pretty much on the whole cat image. So, it will be doing something like this. It's It's going to attend the whole cat image. And that means that those are some So, because those were learned compared to CNNs where we have like the the bias of of locality, here it's obvious that vision transformer found it useful to actually also attend those uh like uh parts of the image that are further away from the query patch, uh which tells us something about uh its usefulness, right? And um so, the bad thing about this whole thing with with vision transformer and transformers in general is that they are they have quadratic complexity. So, that means that uh basically you have n squared complexity, where n is the number of tokens. And here that's the number of of patches in the image, and that's expensive. So, uh basically what this novel uh focal self-attention does is the following. So, the idea is um let's attend to uh nearby patches with a fine-grained structure. So, as you can see here, we'll attend these closer patches with a fine resolution. But once we start getting like further and further away from the query patch, let's kind of start uh just attending to these coarser representations. So, what they'll end up doing is they'll uh they'll basically take these four patches here, and they'll summarize those into this one big patch. And the same thing here. So, here we'll have like I don't know, like maybe 4x4 patches will be summarized into a single patch. And so, that means that we both keep uh the local as well as the global uh like uh connections, and we also reduce the computation, obviously, because we we'll have less less amount of of tokens to attend to. So, that's the main idea idea of this of this of this paper, basically. And as you can see here, so um we're we're going from fine to coarse structure as we're going from closer closer patches towards more like the patches that are further away from from the query patch. Okay, that's that's pretty much it. So now let's start digging into nitty-gritty details and how this everything works like and let's see some results. Okay. So just reiterating here in this paper we represent focal self-attention a new mechanism that incorporates both fine-grained local and coarse-grained global interactions. In this new mechanism each token attends its closest surrounding tokens at fine-grained granularity and the tokens far away at coarse granularity and thus can capture both short and long range short and long range visual dependencies efficiently and effectively. They additionally report that they achieve state-of-the-art results on a couple of different vision benchmarks such as image classification, object detection, and semantic segmentation. And they report some numbers here as well as here. So this is map metric for the object classification object detection, sorry. And yeah, I'll get back to these results a bit later and tell you more about what they what I have to say about those results. Okay, let's now start investigating the high-level architecture. So if we if you focus here the the main difference you'll notice between this architecture and the vision transformer are these parts here. So these denominators. So as you as I told you we we first in order to to kind of transform in order to parse this image by a transformer we'll need to add these patches, right? So we have this patch embedding layer. So the difference between vision transformer and this focal transformer is the following. So basically what they'll do in order to create these patches is in this example what they did is they took four neighboring pixels and that's one patch. So it will have 4 by 4 will be a single patch and that's why you see four here. So because height and width just represent how many pixels you have like in height and width dimensions obviously. Divide by four and you get the number of tokens and D is the dimension of that token after we embed it using this like linear or I think they're using convolutional like projection layer. So, now the the big difference between focal transformer and the like standard vision transformer, so there are two differences. First is they're using this focal self-attention instead of your standard multi-head self-attention. And the second thing is, as you can see, they're reducing the number of these of these like number of of patches, number of tokens. So, as we go into this after you apply this first layer and one times, then what they do is they reduce the number of patches here. So, basically, instead of having this one patch, what you'll have end up having is you'll have this one this will be like a single token, okay? And then going to the next layers here, we'll have even bigger like even bigger tokens, okay? So, one token will basically span this amount of an image, okay? And they are also, as you can notice here, so D goes to 2D, goes to 4D, goes to 8D, and that's a very similar logic that is applied in all convolutional neural networks. So, the idea is as you're reducing the spatial extent of your of your like representation, you're you're increasing the number of channels. And yeah, basically, what that means is we're we're seeing like a huge merge between these transformer and CNN architectures, and it's not even clear what's a transformer what what transformer means and what a CNN means. We also saw those MLP mixers. So, yeah, there is a lot of debate what's convolutional neural network now and what's a transformer actually. Okay. So, that's the that's the high-level overview of how this works. Everything else remains the same as the vision transformer. So, we just have this novel layer, and we have this novel logic where we are reducing the number of of of tokens as we go deeper and deeper. What I show here, this small diagram, let me just tell you briefly about it. Basically, as we are so for a given number of tokens, they show that they show that the receptive field of this novel layer is higher, so it's higher compared to the regular self-attention. And the reason is as we saw is that we are attending as we are getting further away from the query token, we're basically attending coarser and coarser tokens. So we can attend more tokens and have a like a for the same amount of tokens, we can attend like a huge we can have a huge receptive a field. Okay. Now, let me focus on on the actual meat of this paper. And that's this focal self-attention layer. So let me kind of digest how this thing goes together and how they implement this focal self-attention. So first thing you'll notice is so imagine this is your input image and these are just patches. So we have 20 * 20 patches. So that's 20 * 20 tokens. So the first thing they do is this blue blob you can see here are our query tokens. And what they do is they group them into this window. And every every query inside that window will attend to the same group of value and key and keys. And let me just kind of explain what I mean by that. So imagine if we had if we didn't group them in this window fashion, what we would have to do is the following. So this token's neighborhood would actually be something like maybe I don't know like this. And this token's neighborhood here would be maybe something like this. And that means we'd have like a computational overhead. And so basically because of some computational conveniences, they group these into into a window so that they can share the neighborhood. And now I'll show you what I mean by sharing the neighborhood. So first things first, we have these three levels. So level one, level two, level three. And those correspond to those fine-grained over coarse-grained structures. So let me just go back here. So level one was these these fine-grained like patches and then we have these more coarse grained, and finally these are the the coarsest grained, uh, patches, okay? And let me go back here. So, um, the first thing you'll notice is, so we're going to pull these tokens into coarser, uh, representations, as I said. So, that means that here, in level two, we're going to take this two by two, uh, these two by two patches, and we're going to convert them into a single one. So, that's the reason we have, so one, two, three, four, five, six. That's why we have six by six here after after doing the the this this this coarsening. And finally here, because we're on the level three, we're going to we're going to convert these four by four patches into a single into single representation here. So, that's why we have five by five. So, now you can see that on this level one, what we're doing is we are attending this neighborhood here, okay? So, something like this. Let me just draw it. And then on the second level, as you can see here, we're attending this neighborhood, but we have like coarsened representations, okay? And finally, the last layer, let me just draw this, okay? The last layer will attend the whole feature map. And but but the representations will be much coarser. So, that's the idea. And out of those, what we do is we kind of flatten those, and we concatenate those, and then we apply these, uh, like, uh, value uh, we apply a value projection matrices and and key projection matrices to get values and keys. And so, now let let me show you what I mean by this windowed uh, queries. Okay, as you can see here, this same query attends to all of these keys and values, as well as this one. So, all of the queries will be attending to the same keys and values, and that's the whole point here. So, that means, uh, that that's how they kind of, uh, save some computation. So, otherwise, as I already explained, if you had if you had them separated, if you had only this one, then let me just delete this for a second, okay. Whoops, let me just delete all of these and let's focus on a single token, okay? So, if we had only this token, that means that its neighborhood would actually be something like this and then we'd had the coarser one, etc. Whereas for this one here, we'd have a totally different neighborhood. And the fact that these have different neighborhoods means that you'll you'd have to extract all of these for all specific tokens and that's make it more that makes it more computationally intensive to do this and that's why they kind of bucket them into this this window. Okay. That's pretty much it. Then once you have the queries, keys, and values, you do your your regular like self-attention stuff. You basically out of queries and keys, you form those raw scores, then you apply softmax, you get the attention coefficients, and you use those to aggregate value vectors to form novel representations of these tokens here. So, again, these tokens here will have novel representations formed after we apply this layer once. Okay. That's pretty much it. Hopefully that was understandable. Let me now jump further and kind of focus on this part here. So, they say to perform focal self-attention, we need to first extract the surrounding tokens for each query token in the feature map. So, that's the the thing I just mentioned. That's the the motivation behind a kind of grouping these query tokens into into these windows. And I think the Swin Transformer or some of the other like previous work showed that that kind of helps and like this paper is just copying that part. So, note that the strict version of focal self-attention following figure one requires to exclude the overlapping regions across different levels. In our model, we intentionally keep them in order to capture the pyramid information for the overlap regions. So, what they say there is the following thing. So, let me get back back to the image here. As you can see, here looking at this image, it seems like they are they don't have any overlaps. So, that that basically means that they only apply the fine-grained structure here and then for this coarser grain structure, it seems like they're only attending these this region here. So, let me just kind of shade it. So, this part here, okay? But, in reality, they are also they are also using these regions as well for that coarser scale. And yeah, they just mentioned that they they do that and I I I didn't see any ablations on this, but like um it seems that this overlapping information on multiple scales help and they kind of I guess it's also more computationally intensive if they wanted to exclude these regions. So, they they just left it like this. Okay, that's the part. Additionally, they have this this learnable relative position bias that was introduced in this previous previous paper called Swin Transformer and yeah, this bias term kind of just helps boost the performance additionally, okay? Let me just focus on this one now. This is the whole point they have less computation. So, M times N is the M times N are your like feature feature map dimensions. So, that means this is basically the number of tokens N. And as we can see, so the complexity depends on N and we have some like pretty much constant here. So, L is the number of levels and that was three in the examples we just saw. And they also have these S sub R L components. So, let me just kind of show you what these are and as you can see here, so that S sub R, so it's eight for this first level and we can see why it's eight. So, basically that that that's basically the dimension of this pulled feature map, okay? We have 6 by 6 here and that's why we have six here and we have five here, that's why we have five by five here. So, those are pretty much constants and the whole point is that this part is constant once you set up the parameters of your architecture and then you only depend linearly on the number of tokens. Though, this constant may be big, so yeah, that's also worth mentioning and we'll see soon like the the number of flops they have and the time it takes to to to do an inference through this to this architecture. Okay. Uh that's pretty much it. Let me now now focus on on on the actual results. Let's see what they got. So, they have again they have this Focal Tiny, Focal Small, and Focal Base. They tested three of these models with So, this one is obviously the biggest one and they have smaller ones. And um what they report are some like state-of-the-art results, but as we'll see it's not that clear-cut. So, if we focus on this number here, so it's 83.8 and it's a bit bigger than than Swin Base, uh but the thing is they also have more parameters. It's not a significant amount, but it's more. And they also have more flops. And as they show in the appendix as well, it takes more time for the same amount of flops it takes more time to compute uh the forward prop through the Focal Base compared to the Swin Base. So, that's something to keep in mind. If they were to kind of uh make those the same, I'm not sure they'd have even this big of a like a increasing performance. So, I'd say this is more on par than state-of-the-art if you ask me, but like yeah. Uh yeah, if if you if you don't care about these other parameters, then yep, you have a bit better performance. And um they have the same results. They show the same results across different tasks. So, this is the on the COCO. So, this is uh basically object detection. They show some improvements. Again, I'm not sure the underlying flops and and and like number of parameters as well as the inference speed. If you don't care about those again, they do achieve state-of-the-art results there as well. Um Similarly here, so a bunch of different uh bunch of different benchmarks. I won't be focusing on each of those in particular. So, uh Here they show results on semantic segmentation. Again, they're a bit better compared to the baselines. Uh and yeah, it's kind of hard to compare these because they don't have numbers for the number of parameters and and flops there. But yeah, um yeah, I don't want to focus too much on the results. Let me show you some ablations here. Uh in the final model they were actually testing, they used two levels not three levels as we saw in those previous diagrams. And what they tested is increasing the this window size from 7 to 14, and uh that kind of bumped up the the the performance. As you can see here, but also the flops go up. So that's obviously a trade-off again. Uh and what this show here is the the importance of using both the the those local and global like interactions. So first things first, what they show is this uh window attention, which basically means the following thing. So window means that you're If I go back to the diagram up there, so give me a second. So okay, this part. So they that means that these queries are actually these queries are are only attending to the keys and values that are inside of this region, okay? Let me just change the color. So they're only attending this part. They're not attending the other keys and values, and that basically means we don't have a nice communication across different parts of the image. And then obviously uh the performance drops by a lot. Let me get back here, okay? And once they add this uh like local part and then global part, you can see it increases and local plus global, so both using the fine-grained structures, so that's the local part, and the global one, so that's where you're using those windows of size uh 7 or 14, whatever. Uh so when you combine both of those, you get the like computational efficiency plus you get the to to kind of ingest information from the whole image. That's something that we know is useful because that's why transformers are performing better than CNNs. Um okay. Uh what they additionally did here is this Swin Transformer had something like something called like window shift, and it actually uh hurts the performance of Swin Transformer a lot if they remove that shifting, but in for the for the for this focal transformer it's actually not necessary. So, yeah, they just kind of um remove that. Here are some some depth ablations, nothing super super interesting. They again show they are better than Swin uh even even even when they reduce the number of as you can see here even when they reduce depth of the focal one, they are still even better than the Swin that has like like deeper layer in this particular speech of the architecture, okay? And here again they are they are on par. So, yeah, those are just some additional um claims that they are still the art on on on these different tasks, okay? Uh that's pretty much it. Let me just do two more things here. Okay, so what they say here is "Another observation is that adding long-range tokens can bring more relative improvement for image classification than object detection and vice versa for local tokens. We suspect that dense predictions like object detection more rely on fine-grained local context while image classification favors more the global information." So, this is an interesting statement and it kind of makes sense because if you want to understand what's in the image, if you want to classify the object in the image, you want to kind of understand what's like the the whole the whole context, whereas when you're doing object detection it's kind of enough to kind of understand the the fine-grained uh details. But like there is also uh this is a double-edged sword because this could also mean that the model learns the neural networks learn how to exploit those spurious signals that we we are well aware of that sometimes neural networks kind of exploit those spurious signals like um texture in order to to infer what the object is inside of the image and we don't want that. our neural networks to understand the actual content of the image itself and not those spurious signals, okay? Um second thing I want to I want to mention here is they say here "Though extensive experimental results show that our focal self-attention can significantly boost the performance on both image classification and dense prediction tasks. It does introduce extra computational and memory cost since each query token needs to attend to the coarse and global tokens in addition to the local tokens. So, that's what I something I already stressed a couple of times. I don't think this model significantly boosts upon the prior state of the art. And again, they did acknowledge it here. It you have to trade-off computation and so flops and also memory footprint and speed if you want to get this small and not significant boost in performance if you ask me. Um and going back to the appendix here, they mention somewhere here. So, accordingly, our focal transformer has slower running speed though it has similar flops as Swin Transformer. So, that's very important. Uh this is mainly due to two reasons. We introduced the global coarse-grained attention and introduces the extra computations. That's one. And two, though we conduct our focal attention on the window windows, we still observe that extracting the surrounding tokens around windows and the global tokens across the feature map are time-consuming. So, just keep this in mind. If you want to actually run this in production one day, uh those extra parameter those extra dimensions like computation and speed and memory do matter a lot. So, that's pretty much it. Uh hopefully, you like this paper. If you did, share it out, subscribe, and until next time. Bye-bye.

Original Description

❤️ Become The AI Epiphany Patreon ❤️ ► https://www.patreon.com/theaiepiphany In this video I cover a new paper coming from Microsoft: "Focal Self-attention for Local-Global Interactions in Vision Transformers" where they introduce a new transformer layer called focal attention. The main idea is to reduce the complexity but preserve the long-range dependencies. They achieve this by attending to the nearby tokens in a fine-grained manner and to the tokens that are further away they attend their coarsened representations. ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ ✅ Paper: https://arxiv.org/abs/2107.00641 ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ ⌚️ Timetable: 00:00 Main idea of the paper: focal self-attention 04:55 Overview of Focal Transformer architecture 08:15 Focal Self-Attention layer 12:30 Computational complexity, overlapping regions 15:30 SOTA results but with a disclaimer 17:30 Ablations 19:50 Outro, Focal Transformer is slower than Swin ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ 💰 BECOME A PATREON OF THE AI EPIPHANY ❤️ If these videos, GitHub projects, and blogs help you, consider helping me out by supporting me on Patreon! The AI Epiphany ► https://www.patreon.com/theaiepiphany One-time donation: https://www.paypal.com/paypalme/theaiepiphany Much love! ❤️ Huge thank you to these AI Epiphany patreons: Eli Mahler Petar Veličković Zvonimir Sabljic ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ 💡 The AI Epiphany is a channel dedicated to simplifying the field of AI using creative visualizations and in general, a stronger focus on geometrical and visual intuition, rather than the algebraic and numerical "intuition". ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ 👋 CONNECT WITH ME ON SOCIAL LinkedIn ► https://www.linkedin.com/in/aleksagordic/ Twitter ► https://twitter.com/gordic_aleksa Instagram ► https://www.instagram.com/aiepiphany/ Facebook ► https://www.facebook.com/aiepiphany/ 👨‍👩‍👧‍👦 JOIN OUR DISCORD COMMUNITY: Discord ► https://discord.gg/peBrCpheKE 📢 SUBSCRIBE TO MY MONTHLY AI NEWSLETTER: Substack ► https://aie
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Aleksa Gordić - The AI Epiphany · Aleksa Gordić - The AI Epiphany · 52 of 60

1 Intro | Neural Style Transfer #1
Intro | Neural Style Transfer #1
Aleksa Gordić - The AI Epiphany
2 Basic Theory | Neural Style Transfer #2
Basic Theory | Neural Style Transfer #2
Aleksa Gordić - The AI Epiphany
3 Optimization method | Neural Style Transfer #3
Optimization method | Neural Style Transfer #3
Aleksa Gordić - The AI Epiphany
4 Advanced Theory | Neural Style Transfer #4
Advanced Theory | Neural Style Transfer #4
Aleksa Gordić - The AI Epiphany
5 Anyone can make deepfakes now!
Anyone can make deepfakes now!
Aleksa Gordić - The AI Epiphany
6 What is Computer Vision? | The Art of Creating Seeing Machines
What is Computer Vision? | The Art of Creating Seeing Machines
Aleksa Gordić - The AI Epiphany
7 Feed-forward method | Neural Style Transfer #5
Feed-forward method | Neural Style Transfer #5
Aleksa Gordić - The AI Epiphany
8 Alan Turing | Computing Machinery and Intelligence
Alan Turing | Computing Machinery and Intelligence
Aleksa Gordić - The AI Epiphany
9 Feed-forward method (training) | Neural Style Transfer #6
Feed-forward method (training) | Neural Style Transfer #6
Aleksa Gordić - The AI Epiphany
10 What is Google Deep Dream? (Basic Theory) | Deep Dream Series #1
What is Google Deep Dream? (Basic Theory) | Deep Dream Series #1
Aleksa Gordić - The AI Epiphany
11 Semantic Segmentation in PyTorch | Neural Style Transfer #7
Semantic Segmentation in PyTorch | Neural Style Transfer #7
Aleksa Gordić - The AI Epiphany
12 How to get started with Machine Learning
How to get started with Machine Learning
Aleksa Gordić - The AI Epiphany
13 How to learn PyTorch? (3 easy steps) | 2021
How to learn PyTorch? (3 easy steps) | 2021
Aleksa Gordić - The AI Epiphany
14 PyTorch or TensorFlow?
PyTorch or TensorFlow?
Aleksa Gordić - The AI Epiphany
15 3 Machine Learning Projects For Beginners (Highly visual) | 2021
3 Machine Learning Projects For Beginners (Highly visual) | 2021
Aleksa Gordić - The AI Epiphany
16 Machine Learning Projects (Intermediate level) | 2021
Machine Learning Projects (Intermediate level) | 2021
Aleksa Gordić - The AI Epiphany
17 Cheapest (0$) Deep Learning Hardware Options | 2021
Cheapest (0$) Deep Learning Hardware Options | 2021
Aleksa Gordić - The AI Epiphany
18 How to learn deep learning? (Transformers Example)
How to learn deep learning? (Transformers Example)
Aleksa Gordić - The AI Epiphany
19 How do transformers work? (Attention is all you need)
How do transformers work? (Attention is all you need)
Aleksa Gordić - The AI Epiphany
20 Developing a deep learning project (case study on transformer)
Developing a deep learning project (case study on transformer)
Aleksa Gordić - The AI Epiphany
21 Vision Transformer (ViT) - An image is worth 16x16 words | Paper Explained
Vision Transformer (ViT) - An image is worth 16x16 words | Paper Explained
Aleksa Gordić - The AI Epiphany
22 GPT-3 - Language Models are Few-Shot Learners | Paper Explained
GPT-3 - Language Models are Few-Shot Learners | Paper Explained
Aleksa Gordić - The AI Epiphany
23 Google DeepMind's AlphaFold 2 explained! (Protein folding, AlphaFold 1, a glimpse into AlphaFold 2)
Google DeepMind's AlphaFold 2 explained! (Protein folding, AlphaFold 1, a glimpse into AlphaFold 2)
Aleksa Gordić - The AI Epiphany
24 Attention Is All You Need (Transformer) | Paper Explained
Attention Is All You Need (Transformer) | Paper Explained
Aleksa Gordić - The AI Epiphany
25 Graph Attention Networks (GAT) | GNN Paper Explained
Graph Attention Networks (GAT) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
26 Graph Convolutional Networks (GCN) | GNN Paper Explained
Graph Convolutional Networks (GCN) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
27 Graph SAGE - Inductive Representation Learning on Large Graphs | GNN Paper Explained
Graph SAGE - Inductive Representation Learning on Large Graphs | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
28 PinSage - Graph Convolutional Neural Networks for Web-Scale Recommender Systems | Paper Explained
PinSage - Graph Convolutional Neural Networks for Web-Scale Recommender Systems | Paper Explained
Aleksa Gordić - The AI Epiphany
29 OpenAI CLIP - Connecting Text and Images | Paper Explained
OpenAI CLIP - Connecting Text and Images | Paper Explained
Aleksa Gordić - The AI Epiphany
30 Temporal Graph Networks (TGN) | GNN Paper Explained
Temporal Graph Networks (TGN) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
31 Graph Neural Network Project Update! (I'm coding GAT from scratch)
Graph Neural Network Project Update! (I'm coding GAT from scratch)
Aleksa Gordić - The AI Epiphany
32 Graph Attention Network Project Walkthrough
Graph Attention Network Project Walkthrough
Aleksa Gordić - The AI Epiphany
33 How to get started with Graph ML? (Blog walkthrough)
How to get started with Graph ML? (Blog walkthrough)
Aleksa Gordić - The AI Epiphany
34 DQN - Playing Atari with Deep Reinforcement Learning | RL Paper Explained
DQN - Playing Atari with Deep Reinforcement Learning | RL Paper Explained
Aleksa Gordić - The AI Epiphany
35 AlphaGo - Mastering the game of Go with deep neural networks and tree search | RL Paper Explained
AlphaGo - Mastering the game of Go with deep neural networks and tree search | RL Paper Explained
Aleksa Gordić - The AI Epiphany
36 DeepMind's AlphaGo Zero and AlphaZero | RL paper explained
DeepMind's AlphaGo Zero and AlphaZero | RL paper explained
Aleksa Gordić - The AI Epiphany
37 OpenAI - Solving Rubik's Cube with a Robot Hand | RL paper explained
OpenAI - Solving Rubik's Cube with a Robot Hand | RL paper explained
Aleksa Gordić - The AI Epiphany
38 MuZero - Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model | RL Paper explained
MuZero - Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model | RL Paper explained
Aleksa Gordić - The AI Epiphany
39 EfficientNetV2 - Smaller Models and Faster Training | Paper explained
EfficientNetV2 - Smaller Models and Faster Training | Paper explained
Aleksa Gordić - The AI Epiphany
40 Implementing DeepMind's DQN from scratch! | Project Update
Implementing DeepMind's DQN from scratch! | Project Update
Aleksa Gordić - The AI Epiphany
41 MLP-Mixer: An all-MLP Architecture for Vision | Paper explained
MLP-Mixer: An all-MLP Architecture for Vision | Paper explained
Aleksa Gordić - The AI Epiphany
42 DeepMind's Android RL Environment - AndroidEnv
DeepMind's Android RL Environment - AndroidEnv
Aleksa Gordić - The AI Epiphany
43 When Vision Transformers Outperform ResNets without Pretraining | Paper Explained
When Vision Transformers Outperform ResNets without Pretraining | Paper Explained
Aleksa Gordić - The AI Epiphany
44 Non-Parametric Transformers | Paper explained
Non-Parametric Transformers | Paper explained
Aleksa Gordić - The AI Epiphany
45 Chip Placement with Deep Reinforcement Learning | Paper Explained
Chip Placement with Deep Reinforcement Learning | Paper Explained
Aleksa Gordić - The AI Epiphany
46 Text Style Brush - Transfer of text aesthetics from a single example | Paper Explained
Text Style Brush - Transfer of text aesthetics from a single example | Paper Explained
Aleksa Gordić - The AI Epiphany
47 Graphormer - Do Transformers Really Perform Bad for Graph Representation? | Paper Explained
Graphormer - Do Transformers Really Perform Bad for Graph Representation? | Paper Explained
Aleksa Gordić - The AI Epiphany
48 GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation | Paper Explained
GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation | Paper Explained
Aleksa Gordić - The AI Epiphany
49 VQ-VAEs: Neural Discrete Representation Learning | Paper + PyTorch Code Explained
VQ-VAEs: Neural Discrete Representation Learning | Paper + PyTorch Code Explained
Aleksa Gordić - The AI Epiphany
50 VQ-GAN: Taming Transformers for High-Resolution Image Synthesis | Paper Explained
VQ-GAN: Taming Transformers for High-Resolution Image Synthesis | Paper Explained
Aleksa Gordić - The AI Epiphany
51 Multimodal Few-Shot Learning with Frozen Language Models | Paper Explained
Multimodal Few-Shot Learning with Frozen Language Models | Paper Explained
Aleksa Gordić - The AI Epiphany
Focal Transformer: Focal Self-attention for Local-Global Interactions in Vision Transformers
Focal Transformer: Focal Self-attention for Local-Global Interactions in Vision Transformers
Aleksa Gordić - The AI Epiphany
53 AudioCLIP: Extending CLIP to Image, Text and Audio | Paper Explained
AudioCLIP: Extending CLIP to Image, Text and Audio | Paper Explained
Aleksa Gordić - The AI Epiphany
54 RMA: Rapid Motor Adaptation for Legged Robots | Paper Explained
RMA: Rapid Motor Adaptation for Legged Robots | Paper Explained
Aleksa Gordić - The AI Epiphany
55 DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained
DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained
Aleksa Gordić - The AI Epiphany
56 DETR: End-to-End Object Detection with Transformers | Paper Explained
DETR: End-to-End Object Detection with Transformers | Paper Explained
Aleksa Gordić - The AI Epiphany
57 DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!
DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!
Aleksa Gordić - The AI Epiphany
58 DeepMind DetCon: Efficient Visual Pretraining with Contrastive Detection | Paper Explained
DeepMind DetCon: Efficient Visual Pretraining with Contrastive Detection | Paper Explained
Aleksa Gordić - The AI Epiphany
59 Do Vision Transformers See Like Convolutional Neural Networks? | Paper Explained
Do Vision Transformers See Like Convolutional Neural Networks? | Paper Explained
Aleksa Gordić - The AI Epiphany
60 Fastformer: Additive Attention Can Be All You Need | Paper Explained
Fastformer: Additive Attention Can Be All You Need | Paper Explained
Aleksa Gordić - The AI Epiphany

Related Reads

📰
The Research Assistant in the Room
Learn how to build a Research Assistant like Omnist in two weeks with a team of one, leveraging AI and ML concepts
Dev.to · Thomas Lee
📰
AI Mastery: Why Learning AI Is One of the Best Skills Today
Learning AI skills can boost productivity and competitiveness in today's digital world
Dev.to AI
📰
Top AI Papers on Hugging Face - 2026-07-22
Explore top AI papers on Hugging Face, including video grounding and code agents, to stay updated on the latest advancements in AI research
Dev.to · Y Hành Nhan
📰
The Sophistication Trap: Why the Smarter AI Technique Keeps Losing
Smaller AI models can outperform larger ones due to overfitting and complexity, and understanding this phenomenon can inform better AI development strategies
Medium · Deep Learning

Chapters (7)

Main idea of the paper: focal self-attention
4:55 Overview of Focal Transformer architecture
8:15 Focal Self-Attention layer
12:30 Computational complexity, overlapping regions
15:30 SOTA results but with a disclaimer
17:30 Ablations
19:50 Outro, Focal Transformer is slower than Swin
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →