pipette
ESEspañol

HRTF Upsampling Across Varying Measurement Configurations with Geometry-Aware Query-Conditioned Aggregation

Xingyu Chen, Hanwen Bi, Sipei Zhao, Fei Ma, Eva Cheng, Ian S. Burnett

PreprintReal-world use

In the authors' words

Personalized head-related transfer functions (HRTFs) are essential for spatial audio rendering, but densely measuring an individual's HRTFs is costly and time-consuming. HRTF upsampling reduces this burden by estimating dense HRTFs from sparse measurements. Recent learning-based methods have achieved promising performance, but many remain tied to predefined measurement configurations. In this work, we propose GeoAtt, a variable-context HRTF upsampling framework that uses a single trained model across varying measurement configurations. GeoAtt performs geometry-aware, query-conditioned spatial aggregation over the available measurements independently at each frequency bin, followed by frequency-domain modeling using Conformer blocks. The relative geometry between the target and measured directions is incorporated as an additive bias in the cross-attention. Experiments on the SONICOM dataset show that a single trained model achieves the lowest log-spectral distortion across all four canonical Listener Acoustic Personalization (LAP) challenge measurement configurations and further generalizes to configurations that are not explicitly included during training.

Main resultThe abstract does not state a limitation.

Appeared: Wednesday, September 23. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: Submitted to IEEE ICASSP 2027