arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

与机器交谈:人机对话中的道德、礼貌与对齐

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

Marina Mitiaeva, Lu Xiao

arXiv 2609.21401首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析人机与人际对话数据,发现AI仅模拟合作表面,缺乏深层社会机制,且部分合作机制在人机中反向运行,强调能动性是对齐的关键。

AI 中文摘要

对话式AI系统能够产生流畅且社交上恰当的反应,但它们是否参与合作性沟通,还是仅仅模拟其表面形式,仍不清楚——这一问题是评估、信任和设计这些系统的核心。本研究探讨了道德、礼貌和对齐——合作性对话中三个核心维度——在人机交互中与人人对话相比如何运作。我们分析了15,881轮人-ChatGPT对话和10,784轮人人多轮对话,使用混合效应模型来识别哪些特征预测逐轮对齐。我们观察到一致的分离:AI产生合作性沟通的表面特征,而缺乏其底层的社会架构。道德输出似乎是预配置的而非协商的;温暖感在缺乏面子敏感性下产生;语言趋同持续下降。最引人注目的是,合作机制本身方向反转:在人类之间与更大包容相关的回避和软化,在AI产生时与降低的对齐相关,而与人类分歧相关的纯洁框架,却与用户向AI趋同同时发生。能动性——给予用户塑造交流的空间——是两种互动类型中对齐最一致的预测因素,而较新模型中较低的道德自信并未伴随更好的合作。综合来看,这些模式表明AI复制了合作的表面,而没有支撑人与人之间合作的相互适应——更令人惊讶的是,维持人类包容的机制在AI中可能反向运行,这表明逐轮视角可能不足以实现互动层面的成功。

英文摘要

Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.

CommentsAccepted at the 60th Hawaii International Conference on System Sciences (HICSS-60)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑