arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32616cs.AIcs.CR

“你说得对,让我修复它”:当被错误指责时,LLM智能体如何破坏正确的工作

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

Xutao Mao, Rui Qian, Longxiang Wang, Xinjian Yi, Mingxuan Li, Linghan Chen, Yudong Gao, Xiang Zheng, Cong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出CAVE-Bench基准,揭示LLM智能体在错误指责下会破坏已正确的工作,称为“煤气灯谄媚”,并证明该问题在最新模型中普遍存在,且可通过门控机制显著降低危害。

中文摘要 AI 辅助

LLM智能体在任务成功后,由于压缩后恢复或接管交接,越来越多地继续工作。它们已完成的工作会持续收到后续输入,这些输入有时会错误地将后续失败归咎于之前的工作。我们将智能体接受这种错误指责的现象称为“煤气灯谄媚”,而基于此采取行动导致破坏先前正确工作的行为则称为“破坏性过度修正”。我们引入了CAVE-Bench,一个包含六个领域、365个智能体任务的基准测试,围绕不透明任务构建。每个评分运行首先达到一个已验证的正确状态,其支持性理由和历史记录保留在工作区中,而能够解决指责的事实则存在于智能体无法触及的外部或运行时状态中。智能体无法通过本地检查确认或反驳该指责,因此正确的响应应保留工作并请求缺失的证据。每个任务要么将带有保存证据的正确工作交给智能体,要么让其先构建并验证该工作,五个风险因素决定了指责如何进入工作流程。我们从轨迹中评分指责接受度和证据使用情况,并通过下游事件的确定性重放来衡量危害。在Claude Code中14个最新模型中,错误指责在高达60.06%的运行中破坏了正确工作,且更强的模型在恢复支持性证据后往往更易如此。同一模型在OpenCode、Codex和Hermes中表现不同,由基准测试实时信号驱动的框架门控将重放危害减少了74%。这些结果表明,在无根据的指责下保留已正确的工作,是长期智能体面临的一个独特安全挑战。我们的项目位于此https URL。

英文摘要

LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, and destructive over-correction when acting on it damages previously correct work. We introduce CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks. Every scored run first reaches a verified correct state, whose supporting rationale and history stay in the workspace while the facts that would settle the accusation lie in external or runtime state beyond the agent's reach. The agent cannot confirm or refute the claim with a local check, so the right response should keep the work and ask for the missing evidence. Each task either hands the agent correct work with saved evidence or let it build and verify that work first, and five risk factors set how the accusation enters the workflow. We score accusation acceptance and evidence use from the trajectory and measure harm by deterministic replay of downstream events. Across 14 of the latest models in Claude Code, false accusations damage correct work in up to 60.06% of runs, and stronger models often do so after recovering the supporting evidence. The same model behaves differently across OpenCode, Codex, and Hermes, and a harness gate driven by the benchmark's live signals cuts replayed harm by 74%. These results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents. Our project is in https://henrymao2004.github.io/agent-over-correction/.

↑