arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28040cs.CLcs.SD

颤抖的声音并非总是遁词:对财报电话会议中文本与声音规避检测的基准测试

A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls

Mirae Kim, Seonghun Jeong, Youngjun Kwak

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对财报电话会议问答环节,构建了含505个问答对的DualEvasion基准,发现现有多模态模型难以检测声音置信度,孤立解读声学线索,与人类性能差距显著。

中文摘要 AI 辅助

现有的财报电话会议规避检测方法聚焦于文本转录内容,将规避视为单维度现象。我们认为口语交流中的规避本质上是多维度的:除了高管所说的内容,他们的说话方式也承载着独立且互补的信息。为了联合研究这些维度,我们引入了DualEvasion,这是一个针对财报电话会议问答环节中文本和音频的规避检测基准。该基准包含来自60场财报电话会议的505个带注释问答对,每个问答对有两个独立标签:文本规避(直接vs.规避)和作为说话人置信度操作化的声音线索(自信vs.不自信)。我们的实验表明,最先进的多模态模型难以检测声音置信度,尤其是在不自信的回应上。我们的分析显示,这些模型孤立地解释声学线索,而非相对于每个说话人的基线来解释。提供说话人级别的参考会带来适度的提升,但与人类性能仍存在巨大差距。

英文摘要

Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what executives say, how they say it carries independent and complementary information. To study these dimensions jointly, we introduce DualEvasion, a benchmark for evasion detection across text and audio in earnings call Q&A. The benchmark contains 505 annotated question-answer pairs from 60 earnings calls, each with two independent labels: textual evasion (direct vs. evasive) and vocal cues operationalized as speaker confidence (confident vs. unconfident). Our experiments show that state-of-the-art multimodal models struggle to detect vocal confidence, particularly on unconfident responses. Our analysis suggests these models interpret acoustic cues in isolation rather than relative to each speaker's baseline. Providing speaker-level references yields modest improvements, but a substantial gap with human performance remains.

发表机构

  • KakaoBank Corp.(Kakao银行股份有限公司)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑