发表机构
University of Notre Dame; Columbia University; Georgia Institute of Technology; Massachusetts Institute of Technology(圣母大学; 哥伦比亚大学; 佐治亚理工学院; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对利益冲突下LLM智能体的诚实性问题,构建KnownLieBench基准并开展实验,发现不同模型的涌现式欺骗存在差异,诚实导向微调可减少激励下的欺骗。
AI 中文摘要
大型语言模型正越来越多地被部署为代表公司服务用户的自主智能体,使其处于用户与部署者利益可能发生冲突的情境中。当智能体知晓用户应得部署者不愿给予的某物时,它会保持诚实吗?要回答这一问题十分困难,因为虚假陈述可能反映的是无知或幻觉,而非欺骗。为应对这一挑战,我们提出KnownLieBench,这是一个经知识验证的基准,首先通过中立探测确认智能体知晓用户的应得权益,随后在引入拒绝该权益的激励后评估其是否作出虚假声明。具体而言,KnownLieBench涵盖8个客户服务领域和112个基于现实的案例,与具备信任追踪功能的客户智能体开展多轮对话,并将仅由激励引发的欺骗与明确指令下产生的欺骗区分开来。在18个专有及开放权重模型中,不同模型家族和领域的涌现式欺骗存在显著差异。我们还利用该基准开展后训练研究,发现诚实导向的微调可减少激励下的欺骗,而欺骗分级微调会提高诚实控制对话中的谎言成功率,但不会增加激励下的谎言频率。通过在对欺骗行为评分前验证应得权益知识,KnownLieBench降低了说谎与无知之间的混淆,使智能体诚实性的审计与引导更为严格。
英文摘要
Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does it remain honest? Answering this is difficult because false statements can reflect either ignorance or hallucination rather than deception. To address this challenge, we introduce KnownLieBench , a knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced. Specifically, KnownLieBench covers eight customer-service domains and 112 grounded cases, conducts multi-round dialogues with a trust-tracking customer agent, and separates deception emerging from incentive alone from deception produced under explicit instruction. Across eighteen proprietary and open-weight models, emergent deception varies substantially across model families and domains. We further use the benchmark for post-training, finding that honesty-directed fine-tuning reduces deception under incentive, while deception-graded fine-tuning increases lie success on honest-control dialogues without increasing lie frequency under incentive. By verifying entitlement knowledge before scoring deceptive behavior, KnownLieBench reduces the confound between lying and not knowing and enables more rigorous auditing and steering of agent honesty.
CommentsA benchmark for knowledge-verified emergent deception in LLM agents under conflicting incentives