arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

盲选策展人:有偏见的评判者如何在自我进化的智能体中悄然阻碍技能淘汰

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

arXiv 2607.07436首次发表:更新:

发表机构

AWS Generative AI Innovation Center; HSBC Holdings Plc., HSBC Technology Center, China(亚马逊云科技生成式人工智能创新中心; 汇丰控股有限公司,汇丰科技中心,中国)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自我进化智能体中,有偏见的评判者对技能淘汰的影响。通过损坏奖励分析等方法发现,“误判通过”偏差会导致技能淘汰机制失效,此为行为安全结果。还提出廉价审计可判断评判者是否越过阈值影响智能体。

AI 中文摘要

自我进化的智能体通过观察不良技能的失败来淘汰它们,但是当评判者无法看到这些失败时会发生什么呢?技能淘汰是一种结构约束,可防止不断增长的库降至无技能基线以下,但其保障假设奖励是无偏的,而对于无参考任务所依赖的大语言模型评判者来说这是错误的。我们表明有偏见的评判者不仅会增加噪声,还会悄然关闭策展人功能。我们通过损坏奖励分析来精确说明这一点,并通过在确定性奖励之上注入损坏来隔离因果通道,在一个无参考报告撰写测试平台上进行带有代码生成交叉检查的行为研究。对称噪声不会影响淘汰,但“误判通过”偏差(失败被误判为通过)会在超过任何数据量都无法跨越的尖锐阈值后禁用基于贡献的淘汰。将真正的淘汰与上限驱逐 churn 区分开来表明这种机制故障是普遍存在的,适用于各个领域和故障率,只有近乎零误判通过的验证器类评分器除外。然而,下游结果取决于具体情况:只有在相同的损坏也使技能合成匮乏的地方,评估质量才会下降,否则保持稳定,所以被禁用的策展人是“沉默的”,不会在任何总体指标中显现出来。贡献是一个行为安全结果,而非性能结果。一种廉价的缺陷注入审计可以在部署前告诉操作员其评判者处于阈值的哪一侧。

英文摘要

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks require. We show that a biased judge does not merely add noise; it \emph{silently switches off the curator}. We make this precise with a corrupted-reward analysis, then a behavioral study on a reference-free report-writing testbed with a code-generation cross-check, injecting corruption on top of a deterministic reward to isolate the causal channel. Symmetric noise leaves retirement intact, but \emph{false-pass} bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold (here a false-pass rate of $0.45$) that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this \emph{mechanism} failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream \emph{outcome}, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is \emph{silent}, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.

CommentsPublished at COLM 2026 Workshop on Agent Behavior

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑