AI 中文总结
提出VoxParity基准,测试语音智能体是否依据音频线索改变决策;在183个场景中仅11/23系统通过,显示多数系统偏向文字而忽略声音关键信息。
AI 中文摘要
语音智能体可以仅凭文字处理几乎所有的来电,但仍会在其行业规则所针对的少数情形上失败。紧急呼叫标准、欺诈指引、无线电用语和脆弱性规则均认识到,来电者的声音听感或周围可闻声响可能改变正确的行动。VoxParity 测试智能体是否依据这些信息行事。在来自14个行业的183个场景中,转录文本保持不变而音频变化(辅导声音、医疗监护仪蜂鸣、无线电检查下的求救信号、药物名称上的噪声、儿童声音下注、惊恐低语),正确的类型化工具调用也随之改变。一个仅文字的零假设测试仅在听到呼叫比仅读取文字的管道更能改变系统行动时给予系统认可。在23个也可基于转录文本运行的系统里,仅11个通过。描述性地看,错误偏向文字:当音频要求保护时,全部28个系统执行常规请求的频率高于在干净呼叫上过度反应(汇总为41%对12%;仅文字管道为58%对15%)。探索性分析显示,领先系统的大多数失误源于它们听到的线索;系统几乎仅在陈述规则的条目上胜过零假设;领先系统推翻所听到的顺从或困惑的频率远高于急性警报;在测试的模型中,描述声音和陈述规则各自能弥补部分不足,但在情绪上留下缺口。
英文摘要
A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.
Comments38 pages, 11 figures, 15 tables. Code, scorer and development-split data at https://github.com/bhavik-mangla/voxparity-bench and https://doi.org/10.5281/zenodo.23008159