All
Articles 173,180Blog Posts 163,830Tech Tutorials 46,199Research Papers 33,880News 21,840
⚡ AI Lessons
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
arXiv:2608.21761v1 Announce Type: new Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extractin
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
arXiv:2608.05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
arXiv:2608.06150v1 Announce Type: new Abstract: Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
arXiv:2608.05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
arXiv:2608.05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
arXiv:2608.06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Depth-Guided Video Object Counting in Crowded Scenes
arXiv:2608.06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all inst
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
arXiv:2608.06240v1 Announce Type: cross Abstract: Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers
arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
arXiv:2608.04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
arXiv:2608.04154v1 Announce Type: cross Abstract: Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because ter
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
arXiv:2608.04589v1 Announce Type: cross Abstract: EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimod
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
arXiv:2608.04655v1 Announce Type: cross Abstract: Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable g
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Towards a satellite image manipulation and deepfake localization benchmark dataset
arXiv:2608.04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
arXiv:2608.05069v1 Announce Type: cross Abstract: Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
arXiv:2608.05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computat
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
ZoomV: Temporal Zoom-in for Efficient Long Video Understanding
arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
arXiv:2511.19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual unde
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
arXiv:2603.06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce v
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Multi-Camera Trajectory Forecasting with Trajectory Tensors
arXiv:2108.04694v2 Announce Type: cross Abstract: We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajecto
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
arXiv:2403.09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP mo
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
arXiv:2608.03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand gene
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1mo ago
Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
arXiv:2608.03407v1 Announce Type: cross Abstract: Road network segmentation from satellite imagery remains challenging due to large geographic variation in road
DeepCamp AI