AI 中文总结
本文针对三种语音生成系统开展属性级语音控制审计,发现目标响应常伴随非目标变化,提出无训练候选选择器VoDER-Cal,提升了目标响应联合成功率并降低非目标偏差。
AI 中文摘要
自然语言描述已成为控制生成语音的灵活界面。现有评估大多仅衡量输出是否匹配提示词,但仅提示词匹配无法揭示预期变化之外的特征是否保持稳定。我们通过对CosyVoice3、VoxCPM2和Fish-Speech-S2三种语音生成系统的受控配对审计,研究了这一区别。评估包含5940个输出,覆盖6个参考说话人、10段文本、3个随机种子和11种条件。通过声学、韵律、内容和说话人测量,我们发现,处于预期目标方向的响应常伴随描述符特定信号级目标集之外的变化。该模式在目标响应超过基线种子变异的输出中依然存在,且伴随变化在不同系统间差异显著。我们进一步提出VoDER-Cal,一种无训练的候选选择器,在保留足够强的目标响应的同时,偏好更小的非目标偏差。三候选池使联合成功率从单样本直接生成时的4.8%提升至所有候选选择策略下的约14%。在匹配的三候选预算内,VoDER-Cal将未见过的非目标偏差从仅目标选择时的0.344降至0.276,并提升了听者评分的保留度。因此,对保留度敏感的评估可补充提示词依从性评估,而候选重排序提供了实用的推理时改进。代码、配置文件和分析脚本可在this https URL获取
英文摘要
Natural-language descriptions have become a flexible interface for controlling generated speech. Existing evaluations largely assess whether an output matches a prompt, but prompt matching alone does not reveal whether characteristics outside the intended change remain stable. We examine this distinction through a controlled paired audit of three speech-generation systems: CosyVoice3, VoxCPM2, and Fish-Speech-S2. The evaluation contains 5,940 outputs spanning six reference speakers, ten texts, three random seeds, and eleven conditions. Using acoustic, prosodic, content, and speaker measurements, we find that responses in the expected target direction are frequently accompanied by changes outside descriptor-specific signal-level target sets. This pattern remains among outputs whose target response exceeds baseline seed variation, and the accompanying changes differ substantially across systems. We further introduce VoDER-Cal, a training-free candidate selector that retains sufficiently strong target responses while favoring smaller off-target deviations. A three-candidate pool raises the joint success rate from 4.8% under single-sample direct generation to approximately 14% for all candidate-selection policies. Within the matched three-candidate budget, VoDER-Cal reduces held-out off-target deviation from 0.344 under target-only selection to 0.276 and improves listener-rated preservation. Preservation-sensitive evaluation therefore complements prompt-adherence evaluation, while candidate reranking offers a practical inference-time improvement. Code, configuration files, and analysis scripts are available at https://github.com/intelland/VoDER