arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17452cs.CL

对话状态的多模态信号有多可靠?来自远程二元协作任务的证据

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

发表机构科尔比学院
查看机构详情
  • Colby College(科尔比学院)

机构由 AI 辅助整理,请以论文原文为准。

Tahiya Chowdhury

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨从多模态行为测量对话状态的特征可靠性,提出三维评估框架用于视频会议二元对话特征评估,发现语言特征预测佳但跨任务通用性差,声学可靠性受说话者身份影响,交互特征是唯一可靠信号,强调相关评估对对话系统特征选择的重要性。

中文摘要 AI 辅助

从多模态行为中测量诸如认知负荷和对话权力等对话状态,需要不仅具有预测性而且在不同任务背景下都可靠的特征。我们提出了一个三维评估框架,用于评估预测准确性、跨任务通用性和重测可靠性,并将其应用于从视频会议平台上的二元对话中提取的交互、声学和语言特征(AVCAffe数据集;53个二元组,9个任务)。结果表明,没有单一特征家族在所有三个维度上都占主导地位。语言特征对认知负荷的预测准确性最高,但在跨任务评估中表现不佳,显示出对特定任务词汇的敏感性。此外,通常被视为特征稳定性证据的声学可靠性,在控制说话者身份后会下降,这证实了标准韵律特征测量的是语音特征而非对话状态。交互特征提供了唯一真正可靠的信号,在说话者归一化后不变。有趣的是,在所有条件下,对权力角色的分类都接近随机基线,这表明任务级聚合行为在预测对话中权力角色方面存在局限性。我们的发现揭示了三个见解:(1)语言特征预测效果最佳,但在不同任务背景下的通用性较差;(2)一旦控制说话者身份,声学可靠性会降至接近零,这对标准评估实践提出了挑战;(3)交互特征提供了唯一真正可靠的信号,底层优势预测二元组内的认知负荷不对称。这些结果表明,说话者归一化和多维度评估是对话系统中上下文感知、稳健多模态特征选择的先决条件。

英文摘要

Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.

补充信息

↑