pipette
ESEspañol

Self-Supervised Perceptually Interpretable Monocular Depth Estimation

Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis

Preprint with a published version

In the authors' words

Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confidence in safety-critical settings. This paper presents a self-supervised framework for perceptually interpretable monocular depth estimation (PIMDE), designed to associate depth predictions with distinct perceptual components of the input image. Rather than operating directly on RGB inputs, the proposed method decomposes each image into a set of perceptual feature maps (PFMs), each encoding a specific visual cue. Distinct depth estimation branches process these PFMs independently to produce depth estimates (PIDEs), which are subsequently combined through an explicit fusion strategy. This formulation allows us to examine directly the contribution of each perceptual cue to the final depth prediction. Experiments conducted on the KITTI benchmark dataset demonstrate that PIMDE achieves performance comparable to established self-supervised MDE methods while providing additional insight into how different perceptual cues influence depth estimation. These results indicate that perceptual decomposition can support interpretability without sacrificing depth estimation accuracy.

Main resultThe abstract does not state a limitation.

Appeared: Monday, September 28. arXiv. Preprint with a published version.

DOI: 10.1109/ICIP61757.2026.11630094

Published version: 2026 IEEE International Conference on Image Processing (ICIP), pp. 1-6, 2026

Authors' comment: Published at IEEE ICIP 2026; 6 pages, 4 figures