arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越谄媚:大语言模型道德推理中的结构化抵抗与顺从

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

Baihui Wang, Bernard Koch

arXiv 2607.21558首次发表:更新:

AI 中文总结

研究大语言模型道德推理中超越谄媚的结构化抵抗与顺从,通过三项研究揭示其判断修正沿与人类社会心理学现象平行的三个维度结构化,为区分建设性信念修正与谄媚顺从提供原则基础,助力道德交互更好对齐。

AI 中文摘要

构建能向他人学习而不盲目顺从的社会校准大语言模型,仅减少谄媚这一单维度失败模式是不够的。模型必须区分何时纳入他人观点,何时保持有充分依据的道德判断。我们研究了支配这种区分的更广泛的抵抗 - 顺从过程。通过三项研究表明,模型的判断修正沿三个维度结构化,这与人类社会心理学中的经典现象平行:传入观点与模型初始立场的距离、该观点的来源归因以及支持它的联盟结构。模型通常更易接受相近立场,受呈现为自身先前判断的观点影响更大,对群体压力反应不同。这些发现将谄媚重塑为受社会影响的更广泛判断更新过程的一种表现。我们的框架为区分建设性信念修正和谄媚顺从提供了原则基础,从而在道德上重要的交互中支持更好的对齐。

英文摘要

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence. Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑