Papers nuevos sobre Visión por computadora
288 papers nuevos sobre visión por computadora en los últimos 7 días, dentro de IA y aprendizaje automático. Acá están los 50 que Pipette considera más valiosos, con el resultado principal en palabras de sus autores.
Lo mejor de la semana
ZIL: Zero-shot Image-to-LiDAR Registration
We propose ZIL, the first foundation model for zero-shot non-synchronized image-to-LiDAR registration.
PreprintDice ser un gran avanceUso en el mundo realCódigo disponibleAn Eternal Irradiance Camera
We present an omnidirectional irradiance camera that measures the irradiance function---the illumination incident upon every point on a sphere.
Preprint con versión publicadaUso en el mundo realDynamic Thermal Gaussians: Multimodal 4D Gaussian Splatting
To address this limitation, we propose the first dynamic RGB-Thermal reconstruction framework for complex scenes.
PreprintDice ser un gran avanceCódigo disponibleOpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection
We present OpenSAL360, the first open-source platform for scalable, low-cost 360{\deg} video saliency collection.
PreprintUso en el mundo realCódigo disponibleCOVER: Codec-Robust Video Watermarking with Generative Video Priors
We present COVER, the first learned video watermark built around codec compression as its design target, which survives that compression by embedding the payload in the latent space of a frozen generative video autoencoder and recovering it by re-encoding the received video into that same latent space.
PreprintUso en el mundo realnnFoundation: 3D Foundation Models for Radiology
Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging.
PreprintUso en el mundo realCan Spiking Neural Networks play pinball? A neuromorphic motion detector for target tracking
This work presents a fully spiking, real-time perception-to-action pipeline for closed-loop pinball gameplay.
PreprintTrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceBrainIAC: Interactive 3D Brain Lesion Segmentation across Heterogeneous MRI Modalities with Online Adaptation
We present BrainIAC (Brain lesion Interactive Adaptive Continuously learning segmentation), a unified framework that integrates (i) a multi-modal backbone network trained to segment multiple types of brain lesions and handle heterogeneous sets of modalities via zero-filling and random modality dropping; (ii) 3D interactive segmentation with bounding-box and click prompts that preserves fully automatic prediction when no prompt is given; and (iii) an online adaptation mechanism combining Mid-Interaction adaptation and Post-Interaction adaptation, supervised by the network's own predictions as pseudo labels and guided by an extra Click-Centered Gaussian loss.
PreprintUso en el mundo realCódigo disponibleAnnual Earth-observation embeddings encode wildfire disturbance and support simplified burned area mapping
Annual embeddings nevertheless achieve high segmentation accuracy while moving the burden of dense time series processing upstream, providing a promising path towards simpler regional burned area mapping.
PreprintUso en el mundo realA Scene Language Model for Open-Vocabulary Scene Mapping
These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation.
PreprintUso en el mundo realBronchoTop: Bronchoscopy Navigation via RGB-Only Topological Localization
This work presents BronchoTop, a real-time, RGB-only framework for topological bronchoscopy localization that eliminates the need for patient-specific data.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realEnhancing Shrimp Disease Detection via Deep Learning and Data Refinement for Resilient Aquaculture
This work contributes the first application of Vision Transformers (ViT) and Self-Supervised Learning (SSL) to the shrimp farming domain, addressing both performance bottlenecks and data labeling challenges.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realSemi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms
We present a semi-automated workflow that converts a single 360-degree video of a room into individually editable, simulation-ready geometry assets.
PreprintUso en el mundo realLearned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video
We instead learn a model that predicts this transformation in a single forward pass: a MobileNetV4 backbone with FiLM-based emotion conditioning outputs parameters for differentiable global transformations.
PreprintUso en el mundo realToward a foundation model for forest point clouds
These findings identify the practical regime in which pretrained representations are most valuable and suggest that instance discrimination, rather than forest semantics, is the main remaining obstacle to a general-purpose 3D forest foundation model.
PreprintUso en el mundo realCódigo disponibleCube-Splat: High-Fidelity 360{\deg} Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360{\deg} frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center.
PreprintCódigo disponibleSpatiotemporal Flux Probing for Single-Photon Videography
We demonstrate that our approach (1) recovers fast motion and temporal illumination dynamics with substantially fewer photons than prior methods, (2) enables velocity-selective videography that automatically refocuses video onto specific detected motions, and (3) generalizes across sensing modalities including single-photon, event, and spike cameras.
Preprint con versión publicadaScout: Open-World Species Recognition on the Edge
We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model.
PreprintUso en el mundo realPartLLM: A Unified Multimodal Foundation for 3D Part Segmentation
To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition.
PreprintWildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.
PreprintPrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning
We present PrismGPT, a Vision-Language Model (VLM) framework that produces structured, region-aware editing plans from a single input image without relying on commercial black-box tools.
Preprint con versión publicadaUso en el mundo realMonitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning
Using a DAS deployment along a roadway network in the City of College Station, Texas, USA, this study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed.
PreprintUso en el mundo realIndustrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space.
PreprintUso en el mundo realCódigo disponibleHeartian: Physiology-Aware Relightable Gaussian Head Avatar
We propose Heartian, a physiology-aware modulation framework that learns cardiac-cycle-dependent per-frame albedo modulation of facial skin-region Gaussians within a relightable head avatar to encode remote photoplethysmography (rPPG) signals.
Preprint con versión publicadaA Principled Approach to Unsupervised Anomaly Detection
We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation.
PreprintUso en el mundo realCódigo disponibleTraining-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models
We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training.
PreprintUso en el mundo realLatent Commonality Expectation-Maximisation for Box-supervised Tree Crown Instance Segmentation
By leveraging frozen self-supervised features, LACE matches or surpasses fully-supervised specialist baselines from boxes alone, removing the need for polygon annotation in tree crown instance segmentation for sparse-canopy environments where labelled data is scarce.
PreprintUso en el mundo realStyle as Cover: Deep Image Steganography via Stylized Transmission
In this paper, we propose StyleStegaNet, a stylized image hiding framework that replaces cover matching with style-concealment transmission.
PreprintVideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph
We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene.
PreprintMECAIL: Communication-Aware Incremental Learning for Object Detection with 14.6 KB Spatiotemporal Experts
We further introduce Mixture-of-Experts for Communication-Aware Incremental Learning (MECAIL), the first method that meets this strict requirement, in which each new domain or environment is served by a small expert network that adapts the base model.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realAdaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering
We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering.
PreprintUso en el mundo realCódigo disponibleTRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection
Motivated by this observation, we propose TRACE (\emph{Trajectory Representation and Consistency Estimation}), a generation-process-aware framework for AI-generated video detection.
PreprintUso en el mundo realOceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning
We propose OceanXL, a fast and scalable 3DGS-based framework for large-scale underwater reconstruction.
PreprintUso en el mundo realCan Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows
Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.
PreprintGaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
We present \textsc{GaitVista}, a reliability-aware measurement layer whose lightweight gate assigns joint- and frame-specific visual contributions using camera coverage, local visual quality, cross-modal disagreement, and root-motion continuity, and exposes them for inspection.
PreprintUso en el mundo realFrom Change Captions to Change Detection: Semantic-Appearance Agreement Framework for Remote Sensing Change Detection
Therefore, we introduce change-caption-guided RSCD, using change captions as the sole task-specific supervision to learn change masks without manually annotated change masks.
PreprintUso en el mundo realCódigo disponiblePETR: Prompt Ensembling with Training-free Routing for Vision-Language Models
To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-prompt architecture: two complementary prompts are learned from different data and objectives to emphasize seen class discrimination and unseen-class generalization, respectively.
PreprintEvidence-gated multimodal parsing and vectorization of architectural floor plans
We introduce SALI-FP, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence.
PreprintUso en el mundo realBreaking Weather-Content Coupling: Type-Severity Guided Progressive Disentanglement for All-in-One Infrared Restoration
To address these issues, we propose TSGPD-IR, a type-severity guided progressive disentanglement network for all-in-one infrared restoration that factorizes restoration guidance into task-level weather semantics and region-level degradation severity.
PreprintUso en el mundo realColon3R: Cross-Domain 3D Reconstruction from Monocular Colonoscopic Video
We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations.
PreprintUso en el mundo realPrivacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB
Experiments on ScanNet show that our method achieves the best 2D and 3D segmentation performance among privacy-preserving approaches and delivers the strongest zero-shot transfer to SUN RGB-D and SceneNN.
PreprintUso en el mundo realSTAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar
To address these limitations, we propose a Scene- and Task-Aware Radar (STAR) Preprocessor together with an end-to-end training framework.
PreprintUso en el mundo realSAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
To leverage strong 2D and 3D priors jointly, we propose SAM-V (Geometry-Aware Segment Anything for Multi-View Instance Segmentation).
PreprintCódigo disponibleSmall yet Assistive: Spatially-Aware Post-Training for Low Vision
We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection.
PreprintUso en el mundo realMulti-viewpoint Geo-localization with Event Cameras
Here, we introduce an event-based visual place recognition (VPR) system that performs robustly under viewpoint changes.
PreprintCódigo disponiblePrompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings
No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity.
PreprintVision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras.
PreprintUso en el mundo realSeeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
These results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.
PreprintPanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline
PanoSeg3R achieves state-of-the-art performance on panoramic 3D semantic segmentation, improving 3D mIoU by up to 16.02 on ScanNet++, while the curated training data further improves zero-shot performance by up to 4.26 and 43.28 mIoU on Stanford2D3D and ToF-360, respectively.
Preprint