AudioLens:基于推理音频-语言模型的多视角语音聚类
AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
浏览论文内容
中文总结 AI 辅助
该研究提出音频多视角聚类任务,构建基准AudioLens-Bench,开发经推理蒸馏和偏好优化训练的AudioLens-R1模型,实验显示其在语音聚类任务中性能优于基线模型。
中文摘要 AI 辅助
语音聚类是组织快速增长的语音集合的基础任务,可支持对话分析、语音驱动发现等应用。但现有方法依赖固定声学相似度度量或基于ASR的文本流水线,限制了其在不同用户指定视角下重组同一语音集合的能力,尤其当聚类同时依赖语言和副语言线索时。我们提出音频多视角聚类,即模型根据自然语言视角直接划分语音记录,同时推断聚类数量及其分配方式。为研究该设置,我们构建了涵盖多个应用领域的基准AudioLens-Bench,评估视角内和跨视角泛化能力。我们进一步提出AudioLens-R1,一种经推理蒸馏和偏好优化训练的端到端大型音频-语言模型。实验表明,AudioLens-R1的表现始终优于所有基线,整体调整兰德指数(ARI)提升12.99个百分点,V测度提升11.62个百分点。这些结果证明,原生音频-语言模型在语音集合上开展灵活的、视角条件化的结构发现具有应用前景。
英文摘要
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.
发表机构
- University of California, Irvine(加利福尼亚大学欧文分校)
- Dartmouth College(达特茅斯学院)
- Purdue University Northwest(普渡大学西北分校)
机构由 AI 辅助整理,请以论文原文为准。