发表机构
University of California San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NSV-Shift是一个对比基准,用于评估语音到语音模型对非言语发声的理解与响应适应能力,通过22对人工验证的对话测试五个模型,发现模型检测NSV优于解释情绪和差异化响应。
AI 中文摘要
我们引入了NSV-Shift,一个对比基准,用于评估语音到语音模型能否理解非言语发声(NSV)并相应地调整其响应。每一对包含两段词汇内容相同的对话,仅在最后一轮中嵌入的NSV上有所不同。我们的试点包含22对经人工验证的对话(44个音频条件),并评估了五个模型在NSV感知、情绪理解和响应适应方面的表现。结果表明,模型在检测NSV方面通常优于解释其细粒度情绪含义或产生适当差异化响应。数据构建流程、数据集和评估流程可在以下网址公开获取:此https URL。
英文摘要
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.
CommentsTechnical Report