跨不同测量配置的HRTF上采样:基于几何感知查询条件聚合
HRTF Upsampling Across Varying Measurement Configurations with Geometry-Aware Query-Conditioned Aggregation
浏览论文内容
中文总结 AI 辅助
针对HRTF上采样依赖固定测量配置的问题,提出GeoAtt框架,利用几何感知查询条件聚合与Conformer频域建模,在SONICOM数据集上以单一模型在多种配置下达到最低对数谱失真并泛化至未见配置。
中文摘要 AI 辅助
个性化头相关传递函数(HRTFs)对于空间音频渲染至关重要,但密集测量个体的HRTFs成本高昂且耗时。HRTF上采样通过从稀疏测量中估计密集HRTFs来减轻这一负担。近期基于学习的方法已取得令人瞩目的性能,但许多方法仍受限于预定义的测量配置。在本工作中,我们提出GeoAtt,一个可变上下文HRTF上采样框架,它使用单个训练模型即可应对不同的测量配置。GeoAtt在每个频率仓上独立地对可用测量执行几何感知的、查询条件的空间聚合,随后使用Conformer块进行频域建模。目标方向与测量方向之间的相对几何关系作为加性偏置被纳入交叉注意力中。在SONICOM数据集上的实验表明,单个训练模型在所有四种规范的Listener Acoustic Personalization(LAP)挑战测量配置中均实现了最低的对数谱失真,并且进一步泛化到训练中未明确包含的配置。
英文摘要
Personalized head-related transfer functions (HRTFs) are essential for spatial audio rendering, but densely measuring an individual's HRTFs is costly and time-consuming. HRTF upsampling reduces this burden by estimating dense HRTFs from sparse measurements. Recent learning-based methods have achieved promising performance, but many remain tied to predefined measurement configurations. In this work, we propose GeoAtt, a variable-context HRTF upsampling framework that uses a single trained model across varying measurement configurations. GeoAtt performs geometry-aware, query-conditioned spatial aggregation over the available measurements independently at each frequency bin, followed by frequency-domain modeling using Conformer blocks. The relative geometry between the target and measured directions is incorporated as an additive bias in the cross-attention. Experiments on the SONICOM dataset show that a single trained model achieves the lowest log-spectral distortion across all four canonical Listener Acoustic Personalization (LAP) challenge measurement configurations and further generalizes to configurations that are not explicitly included during training.