发表机构
NetoAI(NetoAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对作为有状态多轮工具介导系统的客户服务语音AI,提出验证门控审计框架,区分架构、匹配服务事实、多维度验证并记录结果与负担,以评估其偏见与安全性。
AI 中文摘要
语音AI系统越来越多地用于客户服务交互,来电者的口音、情感、流利度和紧急程度等呈现线索与服务请求一同可用。现有的公平性和安全性评估涵盖语音识别差异、口语对话偏见和语音智能体能力,但很少将客户服务语音智能体视为有状态、多轮、工具介导的系统,其中伤害可能表现为在最终拒绝前的额外负担。我们为此类系统形式化了一种验证门控审计框架,该框架:(i)区分原生语音到语音、级联ASR到语言模型到TTS以及混合工具介导的架构;(ii)在受控的来电者呈现条件下使用匹配的服务事实;(iii)在推理前验证事实不变性、呈现线索、人工制品和声学测量;(iv)记录实质性结果和服务负担路径。我们定义了研究问题、方法论、七个验证门、六组指标集以及一个活跃行业评估计划的声明边界,并通过一个完全合成的退款争议审计实例工作示例说明该框架,生产系统结果未包含在本发布中,公共报告受验证协议约束。
英文摘要
Voice AI systems increasingly mediate customer care interactions where caller presentation cues such as accent, affect, fluency, and urgency are available alongside the service request. Existing fairness and safety evaluations cover speech recognition disparities, spoken dialogue bias, and voice agent capability, but rarely treat customer care voice agents as stateful, multi turn, tool mediated systems where harm can appear as additional burden before any final denial occurs. We formalize a validation gated audit framework for such systems. The framework (i) separates native speech to speech, cascaded ASR to language model to TTS, and hybrid tool mediated architectures; (ii) uses matched service facts across controlled caller presentation conditions; (iii) validates fact invariance, presentation cues, artifacts, and acoustic measurements before inference; and (iv) records both material outcomes and path to service burden. We define the research problem, methodology, seven validation gates, a six family metric set, and claim boundaries for an active industry evaluation program. We illustrate the framework with a fully synthetic worked example of a refund dispute audit instance. Production system results are excluded from this release; public reporting is gated by the validation protocol.