All
Articles 180,708Blog Posts 165,066Tech Tutorials 48,353Research Papers 35,695News 22,580
⚡ AI Lessons
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2d ago
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bound
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2d ago
A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products
arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for envi
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2d ago
CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026
arXiv:2609.18206v1 Announce Type: cross Abstract: The MUMU track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge requires a single unified mu
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3d ago
Same Answer, Different Representations: Hidden instability in VLMs
arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implic
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3d ago
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
arXiv:2505.10664v2 Announce Type: replace-cross Abstract: Verifying the authenticity of AI-generated images presents a growing challenge on social media platfor
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
4d ago
MedSAM3: Delving into Segment Anything with Medical Concepts
arXiv:2511.19046v2 Announce Type: replace-cross Abstract: Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizabi
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
4d ago
GP-VM$\times$SMA: Benchmarking General-Purpose Vision Models and Specialized Architectures for 2D Medical Image Segmentation
arXiv:2603.13044v2 Announce Type: replace-cross Abstract: Medical image segmentation (MIS) is a fundamental component of computer-assisted diagnosis and clinica
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
5d ago
Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving
arXiv:2609.12606v1 Announce Type: new Abstract: While multimodal reasoning has advanced rapidly, solving complex geometry problems critically hinges on active v
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
Hyperbolic Geometry for Open-World Object Detection in Remote Sensing Imagery
arXiv:2609.09626v1 Announce Type: cross Abstract: Open-world object detection (OWOD) extends closed-set detection by requiring models to identify unknown object
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
Distilling Image Prototypes for Guided Test-Time Adaptation
arXiv:2609.09737v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) enhances the robustness of models against distribution shifts but faces two critica
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1w ago
Albedo Estimation via Latent Bridge Matching
arXiv:2609.09884v1 Announce Type: cross Abstract: Recent advances in Intrinsic Image Decomposition (IID) have increasingly relied on generative models. However,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation
arXiv:2608.30727v1 Announce Type: cross Abstract: Small-object detection under long-tailed data distributions is a fundamental yet challenging problem in multim
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring
arXiv:2608.31074v1 Announce Type: cross Abstract: We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YO
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
arXiv:2608.27860v1 Announce Type: cross Abstract: Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estim
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection
arXiv:2509.24192v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary obj
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
2w ago
Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation
arXiv:2510.17198v2 Announce Type: replace-cross Abstract: Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infra
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
arXiv:2608.26856v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Q
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
arXiv:2608.27066v1 Announce Type: cross Abstract: Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
arXiv:2608.21761v1 Announce Type: new Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extractin
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows
arXiv:2608.21460v1 Announce Type: cross Abstract: Vision Language Models have recently shown improvements in several objective and verifiable domains such as ob
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging
arXiv:2503.05916v2 Announce Type: replace-cross Abstract: Accurate segmentation of anatomical structures in ultrasound (US) images, particularly small ones, is
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
arXiv:2608.19739v2 Announce Type: replace-cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
3w ago
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct
DeepCamp AI