arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Astar:为自演进工业AI系统学习提出进化方向

Astar: Learning to Propose Evolution Directions for Self-Evolving Industrial AI Systems

Jinxin Hu, Hao Deng, Haibo Xing, Lingyu Mu, Muyu Zou, Weiqin Yang, Sirui Chen, Bohao Wang, Zhezheng Hao, Hao Zhang, Zulong Chen, Shizhun Wang, Yu Zhang, Xiaoyi Zeng, Jiawei Chen

arXiv 2608.27287首次发表:更新:

AI 中文总结

本研究提出Astar,通过工业系统迭代历史训练专用进化引导模型,解决通用LLM的不足,在Lazada广告系统验证中,其提案成功率远超人类专家与通用LLM,实现完全自动迭代并提升广告系统指标。

AI 中文摘要

现代AI系统通过持续迭代实现进步,该迭代流程包括提出进化方向、编写代码、模型训练与评估。尽管后三个阶段的自动化程度不断提升,但作为起点的有效进化方向提出仍是关键瓶颈,目前仍高度依赖资深专家。本研究探索AI是否可承担该角色。研究发现,通用大语言模型(LLM)即便为先进的GPT-5.5,也仅能提供通用且错位的建议:所需专业知识通过经验积累而非明确编码,因此难以直接注入。为此,我们提出Astar,这是一种基于训练的方法,可从工业系统丰富的迭代历史中学习专用的进化引导模型。然而,实现该想法面临四项挑战:稀疏监督、数据噪声、方向空间庞大以及验证成本过高。我们从两方面解决这些问题:在数据层面,设计了一条流水线,通过成对样本扩展和噪声过滤,将含噪声的历史提交转化为大型、干净的进化语料库;在模型层面,通过中间训练(mid-training)、监督微调(SFT)和强化学习(RL)训练模型,利用分层提示引导进化方向生成,并将RL中的奖励模型作为快速替代评估器。Astar已部署于阿里巴巴Lazada广告系统用于进化方向提出。Astar-8B在实际执行评估中,单提案成功率达0.6786,远超人类专家(0.3229)和最强通用LLM(0.3071)。更重要的是,Astar闭合了迭代循环,实现了完全自动迭代:它在两周内引导了20次连续迭代,使离线Hitrate@200提升23.6%,在线A/B测试则带来GMV相对提升4.86%、广告收入相对提升1.82%。

英文摘要

Modern AI systems advance through continuous iteration: a loop of proposing evolution directions, implementing code, training, and evaluation. While the latter three stages are increasingly automated, the starting point --- proposing effective evolution directions --- remains a critical bottleneck that still relies heavily on senior experts. In this work, we explore whether AI can take over this role. We find that general-purpose LLMs, even the advanced GPT-5.5, offer only generic and misaligned suggestions: the required expertise is accumulated through experience rather than explicitly codified, and thus hard to inject directly. To this end, we propose Astar, a training-based approach that learns a specialized evolution-guiding model from the abundant iteration histories of industrial systems. Realizing this idea, however, raises four challenges: sparse supervision, noisy data, a vast direction space, and prohibitively expensive verification. We address them along two fronts. On the data side, we design a pipeline that turns noisy historical commits into a large, clean evolutionary corpus via pairwise sample expansion and noise filtering. On the model side, we train the model through mid-training, SFT, and RL, guiding evolution direction generation with hierarchical hints and using the reward model in RL as a fast surrogate evaluator. Astar has been deployed in Alibaba's Lazada advertising system for evolution direction proposal. Astar-8B achieves a single-proposal success rate of 0.6786 in real-execution evaluation, far exceeding human experts (0.3229) and the strongest general-purpose LLM (0.3071). More importantly, Astar closes the loop and enables fully automatic iteration: it guided 20 consecutive iterations over two weeks, improving offline Hitrate@200 by 23.6%, while an online A/B test yielded relative lifts of 4.86% in GMV and 1.82% in advertising revenue.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑