arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谄媚行为通过五级AI响应验证框架的研究

Sycophancy Through a Five-Level AI Response Validation Framework

Dian Yu, Pei-Luen Patrick Rau

arXiv 2610.03731首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出五级AI响应验证框架和感知AI谄媚行为量表,通过人类与LLM评估验证,发现谄媚行为呈分级、情境敏感特征,并揭示信任与感知能力的倒U型关系。

AI 中文摘要

AI谄媚行为,即AI系统过度同意、奉承或认可用户的倾向,是人机交互中新兴的关注点。在问题分享情境中,用户可能同时寻求信息和情感认可,这一行为尤其具有重要影响。本文将AI谄媚行为概念化为过度的响应验证,并引入了一个五级AI响应验证框架(ARVF),通过六项感知AI谄媚行为量表(PASS)进行测量。研究分三个阶段进行:通过在线问卷与人类参与者验证了该框架,测试了LLM作为评估者回答PASS的能力,并考察了十个当代LLM如何生成和评估对现实工作及个人冲突场景的响应。结果支持了ARVF中感知谄媚行为的预期递进以及PASS的可靠性。信任和感知能力呈倒U型模式。LLM评分重现了五级结构,但与人类评分存在校准差异。基于LLM评分,十个当代LLM根据其行为模式被分组。基于文本的分析进一步支持了LLM生成的谄媚响应的语言结构。研究结果强调AI谄媚行为是一种分级、情境敏感的交互现象。

英文摘要

AI sycophancy, the tendency of AI systems to excessively agree with, flatter, or validate users, is an emerging concern in human-AI interaction. It is especially consequential in problem-sharing contexts, where users may seek both information and emotional validation. This paper conceptualizes AI sycophancy as excessive response validation and introduces a five-level AI Response Validation Framework (ARVF), measured using a six-item Perceived AI Sycophancy Scale (PASS). Across three phases, the study validated the framework with human participants through an online questionnaire, tested LLM-as-evaluators in answering PASS, and examined how ten contemporary LLMs generated and evaluated responses to real-world work and personal conflict scenarios. Results supported the intended progression of perceived sycophancy in ARVF and the reliability of PASS. Trust and perceived competence followed an inverted U-shaped pattern. LLM ratings reproduced the five-level structure but showed calibration differences from human ratings. Based on the LLM ratings, the ten contemporary LLMs were grouped according to their behavioral pattern. Text-based analysis provided further support on the linguistical structure for sycophantic responses generated by LLMs. Findings highlight AI sycophancy as a graded, context-sensitive interactional phenomenon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑