arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DecoEvo:文本空间中求解器与评分标准生成器技能的分数解耦协同进化

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao

arXiv 2607.25675首次发表:更新:

发表机构

Tsinghua University; Qwen Business Unit of Alibaba; University of Chinese Academy of Sciences; Peking University(清华大学; 阿里巴巴的通义业务部; 中国科学院大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本空间优化中现有方法瓶颈,提出DecoEvo方法,通过解耦目标协同进化求解器与评分标准生成器技能,不依赖黄金评分标准,在多基准和大语言模型主干上表现优异,有显著相对增益。

AI 中文摘要

文本空间优化通过编辑外部自然语言工件而非模型权重来调整大语言模型,优化后的工件可检查且模型可视为黑箱。但多数现有文本空间方法固定评估方式,在开放式任务中会成瓶颈。我们引入DecoEvo,在解耦目标下协同进化求解器技能和评分标准生成器技能,优化时不使用黄金评分标准。求解器技能通过标准级反馈更新,评分标准生成器技能通过独立于求解器总分的需求覆盖和响应辨别互补审计修订。在每个基准的官方评估下,DecoEvo在五个基准和三个大语言模型主干上优于所有比较方法,在五个基准平均水平上比SkillOpt有2.8 - 5.0%的相对增益。

英文摘要

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑