发表机构
Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI); Universität des Saarlandes(德国人工智能研究中心(DFKI); 萨尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对难以直接观察的高阶行为构念,本研究以自我慈悲为例,构建多模态窗口化流程融合视频、音频与文本,在51个反思对话会话上验证了概率级融合优于单模态模型。
AI 中文摘要
在人们学习与成长、情绪调节、对挫折进行反思或在艰难对话中保持对他人觉察的过程中,许多最重要的品质并非直接可观察。它们必须从人们随时间推移的说话方式、动作和声音中推断出来,并且它们抵制大多数机器学习流程所围绕的那种清晰标注。我们通过一个在心理学理论中根基深厚但很少被计算建模的案例来研究这一挑战:自我慈悲,即面对自身挫折时以耐心而非严厉的自我批评来回应的倾向。我们考察了它在技术中介的培训环境中进行的结构化反思访谈中如何显现,在这种环境中,人们自然地谈论社会情感上具有挑战性的情境。由于现有数据集未捕捉此类情境中的此类构念,我们使用一种独立的、时间上重叠的、基于既定理论的标注方案,收集并标注了51个反思性对话会话。我们将基础的六成分心理模型整合为一个三类别监督空间,在自我友善和正念与自我批评或不堪重负状态之间取得平衡,并构建了一个可复现的基于窗口的流程,将视频、音频和文本对齐到共享时间线上。将分别基于每种模态训练的单模态模型与一种简单的概率级融合策略进行比较,后者相对于最佳单模态取得了适度但一致的提升。最后,我们讨论了每种模态在何处成功或遇到困难,这对这类构念在反思性言语中实际表达方式有何启示,以及要更有效地对其及类似构念进行建模需要什么。
英文摘要
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one's own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.
Comments8 pages, 6 figures