arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25177cs.SDcs.AI

AudioLens:基于推理音频-语言模型的多视角语音聚类

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

Wenjun Huang, Qiaosong Chu, Tiger Shao, Pengfei Zhang, Yutong Song, Hanning Chen, Yezi Liu, Weiyi Wu, SungHeon Jeong, Ryozo Masukawa, Sanggeon Yun, Yang Ni, Jiang Gui, Mohsen Imani

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出音频多视角聚类任务,构建基准AudioLens-Bench,开发经推理蒸馏和偏好优化训练的AudioLens-R1模型,实验显示其在语音聚类任务中性能优于基线模型。

中文摘要 AI 辅助

语音聚类是组织快速增长的语音集合的基础任务,可支持对话分析、语音驱动发现等应用。但现有方法依赖固定声学相似度度量或基于ASR的文本流水线,限制了其在不同用户指定视角下重组同一语音集合的能力,尤其当聚类同时依赖语言和副语言线索时。我们提出音频多视角聚类,即模型根据自然语言视角直接划分语音记录,同时推断聚类数量及其分配方式。为研究该设置,我们构建了涵盖多个应用领域的基准AudioLens-Bench,评估视角内和跨视角泛化能力。我们进一步提出AudioLens-R1,一种经推理蒸馏和偏好优化训练的端到端大型音频-语言模型。实验表明,AudioLens-R1的表现始终优于所有基线,整体调整兰德指数(ARI)提升12.99个百分点,V测度提升11.62个百分点。这些结果证明,原生音频-语言模型在语音集合上开展灵活的、视角条件化的结构发现具有应用前景。

英文摘要

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

发表机构

  • University of California, Irvine(加利福尼亚大学欧文分校)
  • Dartmouth College(达特茅斯学院)
  • Purdue University Northwest(普渡大学西北分校)

机构由 AI 辅助整理,请以论文原文为准。

↑