New papers on Computer vision
288 new papers on computer vision in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
ZIL: Zero-shot Image-to-LiDAR Registration
We propose ZIL, the first foundation model for zero-shot non-synchronized image-to-LiDAR registration.
PreprintClaims a big stepReal-world useCode availableAn Eternal Irradiance Camera
We present an omnidirectional irradiance camera that measures the irradiance function---the illumination incident upon every point on a sphere.
Preprint with a published versionReal-world useDynamic Thermal Gaussians: Multimodal 4D Gaussian Splatting
To address this limitation, we propose the first dynamic RGB-Thermal reconstruction framework for complex scenes.
PreprintClaims a big stepCode availableOpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection
We present OpenSAL360, the first open-source platform for scalable, low-cost 360{\deg} video saliency collection.
PreprintReal-world useCode availableCOVER: Codec-Robust Video Watermarking with Generative Video Priors
We present COVER, the first learned video watermark built around codec compression as its design target, which survives that compression by embedding the payload in the latent space of a frozen generative video autoencoder and recovering it by re-encoding the received video into that same latent space.
PreprintReal-world usennFoundation: 3D Foundation Models for Radiology
Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging.
PreprintReal-world useCan Spiking Neural Networks play pinball? A neuromorphic motion detector for target tracking
This work presents a fully spiking, real-time perception-to-action pipeline for closed-loop pinball gameplay.
PreprintTrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates.
PreprintBold claims, read criticallyClaims a big stepBrainIAC: Interactive 3D Brain Lesion Segmentation across Heterogeneous MRI Modalities with Online Adaptation
We present BrainIAC (Brain lesion Interactive Adaptive Continuously learning segmentation), a unified framework that integrates (i) a multi-modal backbone network trained to segment multiple types of brain lesions and handle heterogeneous sets of modalities via zero-filling and random modality dropping; (ii) 3D interactive segmentation with bounding-box and click prompts that preserves fully automatic prediction when no prompt is given; and (iii) an online adaptation mechanism combining Mid-Interaction adaptation and Post-Interaction adaptation, supervised by the network's own predictions as pseudo labels and guided by an extra Click-Centered Gaussian loss.
PreprintReal-world useCode availableAnnual Earth-observation embeddings encode wildfire disturbance and support simplified burned area mapping
Annual embeddings nevertheless achieve high segmentation accuracy while moving the burden of dense time series processing upstream, providing a promising path towards simpler regional burned area mapping.
PreprintReal-world useA Scene Language Model for Open-Vocabulary Scene Mapping
These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation.
PreprintReal-world useBronchoTop: Bronchoscopy Navigation via RGB-Only Topological Localization
This work presents BronchoTop, a real-time, RGB-only framework for topological bronchoscopy localization that eliminates the need for patient-specific data.
PreprintBold claims, read criticallyReal-world useEnhancing Shrimp Disease Detection via Deep Learning and Data Refinement for Resilient Aquaculture
This work contributes the first application of Vision Transformers (ViT) and Self-Supervised Learning (SSL) to the shrimp farming domain, addressing both performance bottlenecks and data labeling challenges.
PreprintBold claims, read criticallyReal-world useSemi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms
We present a semi-automated workflow that converts a single 360-degree video of a room into individually editable, simulation-ready geometry assets.
PreprintReal-world useLearned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video
We instead learn a model that predicts this transformation in a single forward pass: a MobileNetV4 backbone with FiLM-based emotion conditioning outputs parameters for differentiable global transformations.
PreprintReal-world useToward a foundation model for forest point clouds
These findings identify the practical regime in which pretrained representations are most valuable and suggest that instance discrimination, rather than forest semantics, is the main remaining obstacle to a general-purpose 3D forest foundation model.
PreprintReal-world useCode availableCube-Splat: High-Fidelity 360{\deg} Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360{\deg} frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center.
PreprintCode availableSpatiotemporal Flux Probing for Single-Photon Videography
We demonstrate that our approach (1) recovers fast motion and temporal illumination dynamics with substantially fewer photons than prior methods, (2) enables velocity-selective videography that automatically refocuses video onto specific detected motions, and (3) generalizes across sensing modalities including single-photon, event, and spike cameras.
Preprint with a published versionScout: Open-World Species Recognition on the Edge
We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model.
PreprintReal-world usePartLLM: A Unified Multimodal Foundation for 3D Part Segmentation
To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition.
PreprintWildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.
PreprintPrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning
We present PrismGPT, a Vision-Language Model (VLM) framework that produces structured, region-aware editing plans from a single input image without relying on commercial black-box tools.
Preprint with a published versionReal-world useMonitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning
Using a DAS deployment along a roadway network in the City of College Station, Texas, USA, this study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed.
PreprintReal-world useIndustrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space.
PreprintReal-world useCode availableHeartian: Physiology-Aware Relightable Gaussian Head Avatar
We propose Heartian, a physiology-aware modulation framework that learns cardiac-cycle-dependent per-frame albedo modulation of facial skin-region Gaussians within a relightable head avatar to encode remote photoplethysmography (rPPG) signals.
Preprint with a published versionA Principled Approach to Unsupervised Anomaly Detection
We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation.
PreprintReal-world useCode availableTraining-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models
We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training.
PreprintReal-world useLatent Commonality Expectation-Maximisation for Box-supervised Tree Crown Instance Segmentation
By leveraging frozen self-supervised features, LACE matches or surpasses fully-supervised specialist baselines from boxes alone, removing the need for polygon annotation in tree crown instance segmentation for sparse-canopy environments where labelled data is scarce.
PreprintReal-world useStyle as Cover: Deep Image Steganography via Stylized Transmission
In this paper, we propose StyleStegaNet, a stylized image hiding framework that replaces cover matching with style-concealment transmission.
PreprintVideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph
We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene.
PreprintMECAIL: Communication-Aware Incremental Learning for Object Detection with 14.6 KB Spatiotemporal Experts
We further introduce Mixture-of-Experts for Communication-Aware Incremental Learning (MECAIL), the first method that meets this strict requirement, in which each new domain or environment is served by a small expert network that adapts the base model.
PreprintBold claims, read criticallyReal-world useAdaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering
We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering.
PreprintReal-world useCode availableTRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection
Motivated by this observation, we propose TRACE (\emph{Trajectory Representation and Consistency Estimation}), a generation-process-aware framework for AI-generated video detection.
PreprintReal-world useOceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning
We propose OceanXL, a fast and scalable 3DGS-based framework for large-scale underwater reconstruction.
PreprintReal-world useCan Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows
Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.
PreprintGaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
We present \textsc{GaitVista}, a reliability-aware measurement layer whose lightweight gate assigns joint- and frame-specific visual contributions using camera coverage, local visual quality, cross-modal disagreement, and root-motion continuity, and exposes them for inspection.
PreprintReal-world useFrom Change Captions to Change Detection: Semantic-Appearance Agreement Framework for Remote Sensing Change Detection
Therefore, we introduce change-caption-guided RSCD, using change captions as the sole task-specific supervision to learn change masks without manually annotated change masks.
PreprintReal-world useCode availablePETR: Prompt Ensembling with Training-free Routing for Vision-Language Models
To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-prompt architecture: two complementary prompts are learned from different data and objectives to emphasize seen class discrimination and unseen-class generalization, respectively.
PreprintEvidence-gated multimodal parsing and vectorization of architectural floor plans
We introduce SALI-FP, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence.
PreprintReal-world useBreaking Weather-Content Coupling: Type-Severity Guided Progressive Disentanglement for All-in-One Infrared Restoration
To address these issues, we propose TSGPD-IR, a type-severity guided progressive disentanglement network for all-in-one infrared restoration that factorizes restoration guidance into task-level weather semantics and region-level degradation severity.
PreprintReal-world useColon3R: Cross-Domain 3D Reconstruction from Monocular Colonoscopic Video
We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations.
PreprintReal-world usePrivacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB
Experiments on ScanNet show that our method achieves the best 2D and 3D segmentation performance among privacy-preserving approaches and delivers the strongest zero-shot transfer to SUN RGB-D and SceneNN.
PreprintReal-world useSTAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar
To address these limitations, we propose a Scene- and Task-Aware Radar (STAR) Preprocessor together with an end-to-end training framework.
PreprintReal-world useSAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
To leverage strong 2D and 3D priors jointly, we propose SAM-V (Geometry-Aware Segment Anything for Multi-View Instance Segmentation).
PreprintCode availableSmall yet Assistive: Spatially-Aware Post-Training for Low Vision
We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection.
PreprintReal-world useMulti-viewpoint Geo-localization with Event Cameras
Here, we introduce an event-based visual place recognition (VPR) system that performs robustly under viewpoint changes.
PreprintCode availablePrompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings
No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity.
PreprintVision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras.
PreprintReal-world useSeeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
These results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.
PreprintPanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline
PanoSeg3R achieves state-of-the-art performance on panoramic 3D semantic segmentation, improving 3D mIoU by up to 16.02 on ScanNet++, while the curated training data further improves zero-shot performance by up to 4.26 and 43.28 mIoU on Stanford2D3D and ToF-360, respectively.
Preprint