ImproveAnyTask:一种用于迭代模型自我改进的自主后训练框架
ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement
浏览论文内容
中文总结 AI 辅助
ImproveAnyTask提出一种自主后训练框架,通过错误归因、策略选择与执行检查迭代改进模型,在11个任务上平均提升18.29/11.97个百分点。
中文摘要 AI 辅助
将通用大语言模型适配到特定任务需要大量人工努力来设计数据和训练策略。持续改进尤其具有挑战性,因为模型更新会改变错误分布,需要不断细化策略。我们引入了ImproveAnyTask,一种自主后训练框架,在有限的计算预算下提高任务性能。该框架从基于梯度的参数优化中汲取灵感,将适配组织为错误归因、更新方向选择和可执行的模型更新。它结合了指标级和案例级分析来识别焦点问题,然后研究有研究支持的策略,并比较其报告的性能提升和复现难度。所选策略被转化为训练数据和训练配置,在全面后训练之前进行小规模执行检查。随后的评估指导模型选择和进一步适配,同时保留经过验证的策略和脚本以供复用。在11个任务中,ImproveAnyTask在Base和Instruct模型上分别实现了平均18.29和11.97个百分点的提升,最大提升为41.96个百分点,在24小时预算内使用相当于8块H20 GPU的资源。
英文摘要
Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from gradient-based parameter optimization, the harness organizes adaptation into error attribution, update-direction selection, and executable model updates. It combines metric-level and case-level analysis to identify a focal problem, then investigates research-backed strategies and compares their reported gains and reproduction difficulty. The selected strategy is translated into training data and a training configuration, with small-scale execution checks preceding full post-training. Subsequent evaluation guides model selection and further adaptation, while validated strategies and scripts are retained for reuse. Across 11 tasks, ImproveAnyTask achieves mean gains of 18.29 and 11.97 percentage points on the Base and Instruct models, respectively, with a maximum gain of 41.96 points, under a 24-hour budget with resources equivalent to eight H20 GPUs.
发表机构
- HKUST(GZ)(香港科技大学(广州))
- Xiaohongshu Inc.(小红书公司)
- NTU(南洋理工大学)
- HKUST(香港科技大学)
- ECNU(华东师范大学)
- ZJU(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。