发表机构
Beijing Language and Culture University; Peking University(北京语言大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TPvG道德决策框架,通过含结果反馈的任务评估发现,LLM道德决策受决策形式和接收者反馈影响,且与人类决策模式存在差异,需评估其高风险交互下的道德稳定性。
AI 中文摘要
现有大语言模型(LLM)的道德评估通常向模型呈现孤立的道德场景并引出一次性决策,忽略了一种已知会深刻影响人类道德行为的因素:结果反馈。我们引入TPvG(基于文本的痛苦与收益对比,Text-based Pain-versus-Gain),该框架改编自人类道德范式,将结果反馈嵌入“不伤害他人与最大化自身收益”的日常道德困境中。TPvG包含五项道德决策任务,从最小上下文的一次性选择逐步过渡到带有明确结果反馈的序列决策。我们的结果表明,LLM的道德决策受决策形式(一次性vs序列)的强烈影响,明确的接收者反馈在不同模型间产生了异质性影响。此外,LLM对明确接收者反馈的响应与人类参考模式存在差异,表明其决策过程可能不同。这些发现强调需要评估LLM的道德行为在高风险交互场景中是否保持稳定。
英文摘要
Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback. We introduce TPvG (Text-based Pain-versus-Gain), adapted from a human moral paradigm, which embeds consequence feedback into an everyday moral dilemma of not harming others versus maximising self-gain. TPvG comprises five moral decision tasks, progressing from minimal-context one-shot choices to sequential decisions with explicit consequence feedback. Our results show that LLM moral decisions were strongly affected by decision format (one-shot versus sequential), and explicit receiver feedback produced heterogeneous effects across models. Furthermore, LLM responses to explicit receiver feedback diverged from the human reference pattern, suggesting potentially different decision processes. These findings highlight the need to evaluate whether LLM moral behavior remains stable in high-stakes interactive settings.
Comments6 pages, 3 figures