arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00652cs.AIcs.LGcs.NE

自我报告并非验证:进化搜索中LLM操作者的环境基础审计

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

Enrong Pan, Ryan Zhou, Ting Hu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对进化式Contexto搜索中的LLM操作者,通过环境基础审计发现其自我报告存在夸大、校准脱节等问题,验证了智能体自我报告不可靠,需与环境核对。

中文摘要 AI 辅助

语言模型智能体越来越多地提出动作、观察外部反馈并解释自身行为,其置信度和理由是便捷的监控信号,但便捷性并非验证。我们引入一种环境基础审计,其中每个中间提议都能获得确切结果。一个语言模型操作进化式Contexto搜索,其反馈函数为每个有效猜测分配确切排名,无需人工标注。在覆盖5种配置和3个模型家族的200次运行中,4种报告配置产生了12249份自我报告。我们测试了三个假设:陈述的置信度是校准的、继承的理由会影响后续提议、基于适应度的选择能提高报告质量。所有三个假设均不成立:操作者夸大了前100名的成功率,幅度达4.8至9.3倍,且校准和区分度在不同模型家族间存在脱节;对754个继承理由的受控干预将真实理由的可测量益处限制在约250个排名范围内;尽管搜索行为差异显著,但基于适应度或随机选择均未产生可检测的选择差异或报告准确性的代际传递。因此,智能体的自我报告应被视为需与环境核对的主张,而非其自身可靠性的证据。

英文摘要

Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded audit in which every intermediate proposal receives an exact outcome. A language model operates an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation. Across 200 runs spanning five configurations and three model families, four reporting configurations produce 12,249 self-reports. We test three assumptions: stated confidence is calibrated, inherited rationales affect later proposals, and fitness-based selection improves report quality. All three fail. Operators overstate top-100 success by factors of 4.8 to 9.3, while calibration and discrimination dissociate across model families. Controlled interventions on 754 inherited rationales bound any measured benefit of the genuine rationale to roughly 250 ranks. Neither fitness-based nor random selection produces a detectable selection differential or parent-to-offspring transmission in report accuracy, despite sharply different search behavior. Agent self-reports should therefore be treated as claims to verify against the environment, not as evidence of their own reliability.

↑