arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39933cs.AIcs.LG

冲突指南:当竞争行为变得可见时,自动研究性能提升

ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible

Binqian Xu, Qiran Zou, Xiangbo Shu, Dianbo Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出冲突指南,通过识别竞争行为并设计探针反馈,在自动研究的两阶段进化中缓解行为冲突,相比仅标量方法,任务和冲突错误分别减少高达28%和14%。

中文摘要 AI 辅助

在设计机器学习模型时,期望的属性往往处于紧张状态:改善一种行为可能会损害另一种行为,因此任务进展可能取决于缓解这种冲突。基于LLM的自动研究系统,通过迭代编辑模型代码并根据标量任务性能反馈保留编辑,在很大程度上忽略了这种权衡。我们发现,标量反馈在搜索早期支持广泛的探索,但它不能揭示编辑如何影响竞争行为。在匹配预算的实验中,随着任务收益减少而引入竞争行为反馈,增加了同时改善两种行为的提案比例,并在标量平台期之外维持进展。为给定模型获取这种反馈需要识别其竞争行为并设计探针来测量它们。为了使竞争行为反馈可操作,我们引入了冲突指南。其可复用的冲突指南技能结合了基于文献的分类法与模型特定的证据,以识别竞争行为并指定探针,供代码代理作为指标实施。进化分两个阶段进行:第一阶段使用任务反馈进行探索;第二阶段使用探针反馈将提案引向冲突缓解,并且仅在探针指示充分缓解时保留边际收益编辑。在五个不同的模型家族中,相对于仅标量的自动研究,冲突指南将任务和冲突相关错误分别减少了高达28%和14%,其收益扩展到其他代码代理。

英文摘要

When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance feedback, have largely ignored this trade-off. We find that scalar feedback supports broad exploration early in search, but it does not reveal how edits affect competing behaviors. In matched-budget experiments, introducing competing-behavior feedback as task gains diminish increases the share of proposals that improve both behaviors and sustains progress beyond scalar-only plateaus. Obtaining this feedback for a given model requires identifying its competing behaviors and designing probes to measure them. To make competing-behavior feedback actionable, we introduce ConflictGuide. Its reusable ConflictGuide-Skill combines a literature-grounded taxonomy with model-specific evidence to identify competing behaviors and specify probes for a code agent to implement as metrics. Evolution proceeds in two stages: Stage I explores with task feedback; Stage II uses probe feedback to steer proposals toward conflict alleviation and retains marginal-gain edits only when probes indicate sufficient alleviation. Across five diverse model families, ConflictGuide reduces task and conflict-related errors by up to 28% and 14%, respectively, relative to scalar-only AutoResearch, with gains extending to other code agents.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Nanjing University of Science and Technology(南京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑