发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低分辨率音频活动识别中训练与部署分辨率不匹配的问题,提出RAST框架,通过保留token级信息和邻域结构压缩高分辨率教师表示并进行局部对齐,在SAMoSA和AudioIMU数据集上提升约7.8%的识别性能。
AI 中文摘要
音频越来越多地用于人类活动识别(HAR),因为它能够捕捉日常环境中的物体交互、环境事件和上下文线索。高分辨率(HR)音频为模型开发提供了丰富的声学信息,但会带来大量的能源和存储成本,并可能暴露敏感的语音内容。低分辨率(LR)音频为部署提供了一种更具隐私保护性和资源效率的替代方案,但降低采样率可能会移除对活动识别至关重要的声学线索,导致显著的性能下降。我们将这种训练-部署不匹配问题形式化为传感器分辨率的特权学习,其中HR音频在训练期间可用,而推理则完全依赖LR音频。我们提出了RAST,一种分辨率感知的迁移框架,通过保留token级信息和邻域结构来压缩HR教师表示,然后进行局部化的HR-LR对齐。在SAMoSA和AudioIMU数据集上的实验表明,RAST始终优于仅使用LR训练和直接教师迁移的基线方法,在仅使用LR音频进行推理的情况下,将仅LR识别性能提升了约7.8%。
英文摘要
Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.