AI 中文总结
本研究提出Astar,通过工业系统迭代历史训练专用进化引导模型,解决通用LLM的不足,在Lazada广告系统验证中,其提案成功率远超人类专家与通用LLM,实现完全自动迭代并提升广告系统指标。
AI 中文摘要
现代AI系统通过持续迭代实现进步,该迭代流程包括提出进化方向、编写代码、模型训练与评估。尽管后三个阶段的自动化程度不断提升,但作为起点的有效进化方向提出仍是关键瓶颈,目前仍高度依赖资深专家。本研究探索AI是否可承担该角色。研究发现,通用大语言模型(LLM)即便为先进的GPT-5.5,也仅能提供通用且错位的建议:所需专业知识通过经验积累而非明确编码,因此难以直接注入。为此,我们提出Astar,这是一种基于训练的方法,可从工业系统丰富的迭代历史中学习专用的进化引导模型。然而,实现该想法面临四项挑战:稀疏监督、数据噪声、方向空间庞大以及验证成本过高。我们从两方面解决这些问题:在数据层面,设计了一条流水线,通过成对样本扩展和噪声过滤,将含噪声的历史提交转化为大型、干净的进化语料库;在模型层面,通过中间训练(mid-training)、监督微调(SFT)和强化学习(RL)训练模型,利用分层提示引导进化方向生成,并将RL中的奖励模型作为快速替代评估器。Astar已部署于阿里巴巴Lazada广告系统用于进化方向提出。Astar-8B在实际执行评估中,单提案成功率达0.6786,远超人类专家(0.3229)和最强通用LLM(0.3071)。更重要的是,Astar闭合了迭代循环,实现了完全自动迭代:它在两周内引导了20次连续迭代,使离线Hitrate@200提升23.6%,在线A/B测试则带来GMV相对提升4.86%、广告收入相对提升1.82%。
英文摘要
Modern AI systems advance through continuous iteration: a loop of proposing evolution directions, implementing code, training, and evaluation. While the latter three stages are increasingly automated, the starting point --- proposing effective evolution directions --- remains a critical bottleneck that still relies heavily on senior experts. In this work, we explore whether AI can take over this role. We find that general-purpose LLMs, even the advanced GPT-5.5, offer only generic and misaligned suggestions: the required expertise is accumulated through experience rather than explicitly codified, and thus hard to inject directly. To this end, we propose Astar, a training-based approach that learns a specialized evolution-guiding model from the abundant iteration histories of industrial systems. Realizing this idea, however, raises four challenges: sparse supervision, noisy data, a vast direction space, and prohibitively expensive verification. We address them along two fronts. On the data side, we design a pipeline that turns noisy historical commits into a large, clean evolutionary corpus via pairwise sample expansion and noise filtering. On the model side, we train the model through mid-training, SFT, and RL, guiding evolution direction generation with hierarchical hints and using the reward model in RL as a fast surrogate evaluator. Astar has been deployed in Alibaba's Lazada advertising system for evolution direction proposal. Astar-8B achieves a single-proposal success rate of 0.6786 in real-execution evaluation, far exceeding human experts (0.3229) and the strongest general-purpose LLM (0.3071). More importantly, Astar closes the loop and enables fully automatic iteration: it guided 20 consecutive iterations over two weeks, improving offline Hitrate@200 by 23.6%, while an online A/B test yielded relative lifts of 4.86% in GMV and 1.82% in advertising revenue.