Foundations
Computer Vision
Object detection, segmentation, YOLO, CLIP, and vision-language models
Skills in this topic
3 skills — Sign in to track your progress
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bound
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products
arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for envi
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1w ago
CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026
arXiv:2609.18206v1 Announce Type: cross Abstract: The MUMU track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge requires a single unified mu
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1w ago
Same Answer, Different Representations: Hidden instability in VLMs
arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implic
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1w ago
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
arXiv:2505.10664v2 Announce Type: replace-cross Abstract: Verifying the authenticity of AI-generated images presents a growing challenge on social media platfor
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1w ago
MedSAM3: Delving into Segment Anything with Medical Concepts
arXiv:2511.19046v2 Announce Type: replace-cross Abstract: Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizabi
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1w ago
GP-VM$\times$SMA: Benchmarking General-Purpose Vision Models and Specialized Architectures for 2D Medical Image Segmentation
arXiv:2603.13044v2 Announce Type: replace-cross Abstract: Medical image segmentation (MIS) is a fundamental component of computer-assisted diagnosis and clinica
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving
arXiv:2609.12606v1 Announce Type: new Abstract: While multimodal reasoning has advanced rapidly, solving complex geometry problems critically hinges on active v
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Hyperbolic Geometry for Open-World Object Detection in Remote Sensing Imagery
arXiv:2609.09626v1 Announce Type: cross Abstract: Open-world object detection (OWOD) extends closed-set detection by requiring models to identify unknown object
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Distilling Image Prototypes for Guided Test-Time Adaptation
arXiv:2609.09737v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) enhances the robustness of models against distribution shifts but faces two critica
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Albedo Estimation via Latent Bridge Matching
arXiv:2609.09884v1 Announce Type: cross Abstract: Recent advances in Intrinsic Image Decomposition (IID) have increasingly relied on generative models. However,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation
arXiv:2608.30727v1 Announce Type: cross Abstract: Small-object detection under long-tailed data distributions is a fundamental yet challenging problem in multim
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring
arXiv:2608.31074v1 Announce Type: cross Abstract: We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YO
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
arXiv:2608.27860v1 Announce Type: cross Abstract: Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estim
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection
arXiv:2509.24192v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary obj
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation
arXiv:2510.17198v2 Announce Type: replace-cross Abstract: Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infra
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
arXiv:2608.26856v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Q
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
arXiv:2608.27066v1 Announce Type: cross Abstract: Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
arXiv:2608.21761v1 Announce Type: new Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extractin
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows
arXiv:2608.21460v1 Announce Type: cross Abstract: Vision Language Models have recently shown improvements in several objective and verifiable domains such as ob
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging
arXiv:2503.05916v2 Announce Type: replace-cross Abstract: Accurate segmentation of anatomical structures in ultrasound (US) images, particularly small ones, is
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
arXiv:2608.19739v2 Announce Type: replace-cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
arXiv:2608.21099v1 Announce Type: cross Abstract: Multi-modal object detection is essential for robust scene understanding in challenging conditions, including
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Breaking the weakest link to evade vision language models
arXiv:2608.18938v1 Announce Type: new Abstract: Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling j
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, c
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
arXiv:2608.16285v1 Announce Type: cross Abstract: Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segme
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes a
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases
arXiv:2608.11582v1 Announce Type: cross Abstract: Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquit
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation
arXiv:2608.11681v1 Announce Type: cross Abstract: This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmen
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers
arXiv:2608.10989v1 Announce Type: cross Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transform
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
arXiv:2608.11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitorin
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Entropy-Centric Explainable AI for Remote Sensing Image Segmentation
arXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. M
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
1mo ago
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
arXiv:2608.08307v1 Announce Type: cross Abstract: Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, bo
Weights & Biases
👁️ Computer Vision
⚡ AI Lesson
1mo ago
When axis-aligned boxes break: Lessons from a CVPR-published traffic AI
A case study in treating representation as a design decision, not a given — with the code and the logging to prove it.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
arXiv:2608.05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
arXiv:2608.06150v1 Announce Type: new Abstract: Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
arXiv:2608.05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
arXiv:2608.05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
arXiv:2608.06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Depth-Guided Video Object Counting in Crowded Scenes
arXiv:2608.06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all inst
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
arXiv:2608.06240v1 Announce Type: cross Abstract: Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired
Reddit r/MachineLearning
👁️ Computer Vision
⚡ AI Lesson
1mo ago
[R], Need some best model suggestions for Face Detection,Face Recognition,Body Detection and Body identification. [R]
need those for analysing movies. example let's say I have to find the screentime of the actor over the whole runtime of the movie and i need to do it for the
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers
arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
arXiv:2608.04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
arXiv:2608.04154v1 Announce Type: cross Abstract: Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because ter
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
arXiv:2608.04589v1 Announce Type: cross Abstract: EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimod
DeepCamp AI