发表机构
Hanyang University; Singapore University of Technology and Design; Singapore Management University(汉阳大学; 新加坡科技设计大学; 新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出跨文化调解基准CC-Mediation及两个DMIS相关评估指标,发现当前大型语言模型在跨文化冲突调解的介入时机与策略上存在局限。
AI 中文摘要
大型语言模型(LLM)开展跨文化调解需同时决定何时介入以及如何回应基于文化的冲突。该问题的进展受限于两点:一是缺乏具备可衡量下游效果的调解数据集,二是缺乏评估跨文化立场转变的原则性指标。为解决这些缺口,我们引入CC-Mediation,这是一个基于跨文化敏感性发展模型(DMIS)的跨文化调解基准,包含1661轮十轮对话,涉及基于文化的冲突、调解介入及介入后轨迹。我们进一步提出两个基于DMIS的评估指标:轨迹AUC,用于衡量跨文化改善随时间的持续性;以及带符号的Wasserstein-1距离,用于衡量跨文化立场转变的幅度与方向。这两个指标与人类对DMIS基础立场转变的判断具有高度一致性。利用CC-Mediation,我们发现当前LLM在两个维度上存在局限:介入时机(何时)的失败源于忽略对话内容的位置先验,而调解策略(如何)的失败则来自后期层的触发崩溃,而非知识缺陷。
英文摘要
Cross-cultural mediation by large language models (LLMs) requires deciding both when to intervene and how to respond in culturally grounded conflicts. Progress on this problem has been limited by the lack of (1) mediation datasets with measurable downstream effects and (2) principled metrics for evaluating intercultural stance change. To address these gaps, we introduce CC-Mediation, a cross-cultural mediation benchmark of $1{,}661$ ten-turn dialogues grounded in the Developmental Model of Intercultural Sensitivity (DMIS), containing culturally grounded conflicts, mediation interventions, and post-intervention trajectories. We further propose two DMIS-based evaluation metrics: Trajectory AUC, which measures the persistence of intercultural improvement over time, and a signed Wasserstein-1 distance, which measures the magnitude and direction of shifts in intercultural stance. Both metrics show strong agreement with human judgment of DMIS-grounded stance shift. Using CC-Mediation, we find that current LLMs have limitations on both axes: intervention timing (when) failure stems from a positional prior that ignores dialogue content, while mediation strategy (how) failure arises from a late-layer elicitation collapse rather than a knowledge deficit.