发表机构
National Research Council Canada(加拿大国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出VoxReason,用于在语音合成前无听者评估基于源的语音规划,通过确定性验证器及实验验证其能有效检测源使用失败,7B模型的SFT+CF修复可提升规划性能。
AI 中文摘要
富有表现力的语音系统在生成任何波形前会做出一个决定:如何传递话语。在对话智能体、旁白和角色条件语音合成(TTS)中,这一隐藏的规划步骤决定了情感、音高、能量、语速、停顿、强调和立场,但下游音频评分很少揭示这些选择是否符合源记录,这种源使用失败发生在任何波形存在之前。VoxReason将这种合成前的决策作为基于源的语音规划的无听者任务进行量化。在合成前,VoxReason会评估传递选择是否基于引用的源记录。系统输出带有证据引用的源引用语音规划,一个确定性验证器会检查引用合法性、槽位一致性、无支撑状态、模式有效性以及单线索反事实局部性。在1440个经过检查的源标记案例中,捷径对照显示仅槽位准确率不可靠:键查找神谕在见过的键上达到1.000的规划槽位准确率,而情感先验在源键不相交的案例中仍达到0.958的槽位准确率,且未引用强度或身份。在另一组100个源键不相交的学习案例比较中,7B局部性SFT+CF修复将规划槽位准确率/局部性从0.684/0.141提升至0.919/1.000,移除源记录会使需要引用的接地分数降低0.488。当前评估不涵盖生成的波形质量。
英文摘要
Plan accuracy alone cannot show whether a speech-delivery decision follows its source: a fixed prior may match the original label yet fail to respond appropriately when a cue changes. VoxReason provides a 100-case verifier benchmark that holds each utterance fixed, edits one designated source-label cue, and scores cited evidence, eight plan fields, and the permitted response. On a source-key-disjoint test of 24 cases, a source-emotion prior reaches plan-slot accuracy 0.958, but none of the 24 edited neutral targets appears in its training labels; its required-change accuracy is 0.000. This diagnoses the support boundary of this prior, not its performance on supported edits. In a complementary 32-case emotion-disjoint test, the prior has seen all edited neutral targets but neither original test emotion; its plan-slot accuracy is 0.219 and required-change accuracy is 1.000. The partitions reuse and overlap the same 100 cases, so these deterministic diagnostics are not independent cohorts or learned-planner results. The benchmark evaluates derived labels and structured plans, not audio input, generated speech, or listener judgments.