多方回馈预测:诊断、基准与上限
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
浏览论文内容
中文总结 AI 辅助
本研究基于AMI语料库构建多方回馈预测基准,发现双人模型零样本失效但特征可重训提升,且回馈预测对未见听者泛化差,揭示身份与线索纠缠,并发布基准与工具。
中文摘要 AI 辅助
回馈预测的研究几乎完全局限于双人对话场景。我们基于AMI语料库引入了一个多方基准,包含来自171次会议、190位说话人的682个掩蔽听者视角,以及18,697个回馈事件,并采用人物不相交的留出划分。一个最先进的双方模型在零样本应用于会议音频时表现如同随机猜测(AUROC为0.499);然而,其冻结的声学特征仍具信息量:线性探针达到0.704,重训练预测器将性能提升至0.751。重训练揭示了第二个局限性。听者条件化改善了训练中见过的听者的预测,但对未见听者无效,且在容量缩减、听者对抗训练、逐听者适应及神谕词汇条件化下,这一差距依然存在。对抗训练仅移除了部分说话人身份信息,而更强的移除反而损害预测,表明身份与对回馈有用的线索纠缠在一起。模型内对照有助于解释这一模式:在相同特征和数据划分下,话轮起始预测可迁移至未见听者,而回馈预测则不能。回馈率在个体间的变异也约为话轮起始率的两倍。由于回馈仅占约1%的帧,帧级F1受基率影响强烈。因此,我们在听者活跃区域报告AUROC及事件F1。我们在该https URL发布了基准和评估工具。
英文摘要
Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.
发表机构
- Enchanted Tools
- IMT Nord Europe
机构由 AI 辅助整理,请以论文原文为准。