StanceBench:一个基于音频大语言模型的语音人际立场评估基准
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
浏览论文内容
中文总结 AI 辅助
研究基于音频大语言模型的语音人际立场评估,引入StanceBench基准,通过Seamless Interaction语料库规范评估,明确9个立场维度,报告大语言模型评判的多项指标,指出不同立场的判断难易及特点。
中文摘要 AI 辅助
语音到语音对话模型越来越依赖韵律和交互细微差别来传达社会意图,但这些线索的基准仍然有限。我们引入了StanceBench,这是一个用于测量对话语音中人际立场并将具备音频能力的大语言模型评估为自动评判者的基准。使用Seamless Interaction语料库,StanceBench(1)通过角色提示极点指定9个立场维度,(2)规范单说话者和基于交互的评估,(3)报告大语言模型作为评判者的稳健性、偏差和立场推理。在评估的立场中,同理心和礼貌最容易判断。温暖和坚定性在具有积极偏差/不对称性的情况下适度可分离。诚实最难判断且显示出高提示顺序偏差,这与需要跨轮证据一致。专注度可分离但与人类的一致性较弱。交互立场对上下文更敏感,存在阈值差距和高方差,尤其是冲突调节方面。
英文摘要
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.
发表机构
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Amazon AGI(亚马逊通用人工智能公司)
- Amazon Research(亚马逊研究院)
机构由 AI 辅助整理,请以论文原文为准。