发表机构
University of Cambridge; University of Oxford; University of Washington Bothell(剑桥大学; 牛津大学; 华盛顿大学博塞尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ParaGeo通过匹配内容分解,在冻结语音模型中建立共享潜在几何,实现副语言属性与内容的解耦,并验证了跨内容泛化能力。
AI 中文摘要
语音表达既随所请求的副语言属性变化,也随语言内容变化。我们引入ParaGeo,一种在冻结的语音语言模型中对副语言变异进行匹配内容分解的方法。合成的音频令牌以固定的监听提示重放;池化的键/值(K/V)表示被居中并投影到一个共享的低维空间。我们的GLM-4-Voice探针跨越12个基准家族的80个请求控制,涵盖八个句子。使用全局拟合的校准基,基于该基的内容保留质心准确率为9.49%,而置换基线为1.25%;同标签跨内容余弦相似度为0.285,对比0.017,两种条件置换检验均得p=0.001。一个独立的十场景、六风格探针揭示了跨场景的可复现对比方向。静态、加性和时间干预产生属性、层和计划依赖的响应轮廓。这些结果为测量副语言结构提供了共享坐标表示,并为潜在语音控制提供了经验起点。代码可在该https URL获取。
英文摘要
Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposition of paralinguistic variation in a frozen speech language model. Synthesized audio tokens are replayed with a fixed listening prompt; pooled key/value (K/V) representations are centered and projected into a shared low-dimensional space. Our GLM-4-Voice probe spans 80 requested controls from 12 benchmark families across eight sentences. With a globally fitted calibration basis, content-held-out centroid accuracy using this basis is 9.49% versus a 1.25% permutation baseline; same-label cross-content cosine similarity is 0.285 versus 0.017, and both conditional permutation tests yield p = 0.001. A separate ten-scenario, six-style probe reveals reproducible contrast directions across scenarios. Static, additive, and temporal interventions produce attribute-, layer-, and schedule-dependent response profiles. These results provide a shared coordinate representation for measuring paralinguistic structure and an empirical starting point for latent speech control. Code is available at https://github.com/yuhanlydia/ParaGeo.