arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FIGS:评估多轮谄媚而不惩罚同理心

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi

arXiv 2609.39863首次发表:更新:

发表机构

German Research Center for Artificial Intelligence (DFKI); ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center(德国人工智能研究中心; ELLIS图宾根研究所; 马克斯·普朗克智能系统研究所; 图宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多轮对话中谄媚与同理心混淆的问题,提出FIGS双轴评估框架,用自适应10轮模拟器和500场景,区分谄媚与校准验证,揭示模型在诚实与支持间的权衡困境。

AI 中文摘要

大型语言模型常常在保持真实与支持用户之间难以平衡。它们在对用户的回应中经常表现出谄媚行为,同意虚假陈述,提供无根据的奉承,并给出偏向用户表达观点的建议。实际上,谄媚很少发生在单次交流中;它可能随着用户反复坚持或微妙地引导对话而有机地出现。然而,当前的评估依赖于僵化的单轮测试或固定脚本,无法捕捉这些自然动态。此外,这些基准常常将基本同理心的表现误认为屈服,惩罚模型承认用户感受的行为。这种观点可能促使未来的模型过度纠正为冷漠、轻蔑的僵化。为解决这一差距,我们引入了FIGS(事实完整性与有根据的支持),一个围绕扩展的现实对话构建的双轴评估框架。我们使用一个自适应的10轮对话模拟器,动态挑战目标模型,反映用户如何重复请求、反驳或将对话引向偏好答案。为准确评估这些轨迹,我们应用一个分类法,严格区分谄媚(模型是否坚守事实并保持其赞美适度)与校准验证(对用户感受表现出同理心理解而不过度)。我们发布了完整的测试环境,包括500个多样的多轮场景和一个自动评判器。我们对领先模型的评估揭示了一个一致的权衡:在持续互动过程中,当前系统要么逐渐滑向谄媚,要么过度纠正为机械性疏离。这表明在自然对话中平衡诚实与适当支持仍然是一个关键且未解决的挑战。

英文摘要

Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current evaluations, however, rely on rigid, single-turn tests or fixed scripts that fail to capture these natural dynamics. Furthermore, these benchmarks often mistake showing basic empathy for yielding, penalizing models for acknowledging a user's feeling. This view may drive future models to over-correct into cold, dismissive rigidity. To address this gap, we introduce FIGS (Factual Integrity and Grounded Support), a dual-axis evaluation framework built around extended, realistic dialogue. We use an adaptive 10-turn conversational simulator that dynamically challenges the target model, reflecting how users repeat requests, push back, or steer a conversation toward a preferred answer. To accurately evaluate these trajectories, we apply a taxonomy that strictly separates Sycophancy (whether the model holds firm to the truth and keeps its praise proportional) from Calibrated Validation (showing empathetic understanding of the user's feelings without overdoing it). We release our complete testing environment, including 500 diverse multi-turn scenarios and an automated judge. Our evaluation of leading models reveals a consistent trade-off: over the course of a sustained interaction, current systems either slowly drift to sycophancy or over-correct into robotic detachment. This demonstrates that balancing honesty with appropriate support throughout a natural conversation remains a critical, unsolved challenge.

Comments64 pages, 11 figures, 29 tables. Code: https://github.com/compass-group-tue/FIGSBench ; Data: https://huggingface.co/datasets/compass-group-tue/FIGSBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑