用于强化学习的从评分规则到代码的信用分配
Rubric-to-Code Credit Assignment for Reinforcement Learning
- Inclusion AI, Ant Group(蚂蚁集团Inclusion AI)
- Zhongnan University(中南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对网页应用代码生成中信用分配被削弱的问题,提出RCCA强化学习框架,构建分层奖励与归因对齐机制,使Ling-RCCA-Flash在MiniAppBench和ArtifactsBench上创出新的最优表现。
AI中文摘要:
交互式网页应用生成需要模型根据自然语言请求生成可用的HTML、CSS和JavaScript应用。与传统代码生成不同,应用质量取决于多个面向用户的功能需求,每个需求通常与局部代码区域相关,如事件处理程序、状态更新、DOM片段或CSS选择器。标准GRPO将这些结构化结果合并为单个序列级奖励,并将所得优势均匀应用于所有token,削弱了信用分配。我们提出从评分规则到代码的信用分配(Rubric-to-Code Credit Assignment,RCCA),这是一种强化学习框架,可将评分规则级别的功能反馈转换为针对生成代码的局部优化信号。RCCA围绕明确的功能评分规则构建训练任务,使用分层奖励区分格式、源代码、运行时和功能故障,并将评估器生成的文本归因与负责的代码跨度及生成的token对齐。最终模型Ling-RCCA-Flash在MiniAppBench上得分为41.25,较Ling-3.0-Flash提升32.20分,且略优于Claude Opus 4.5;在ArtifactsBench上得分为76.19,较SFT模型提升4.48分,在官方ArtifactsBench排行榜设置下超越GPT-5得分3.64分,创下新的最高得分,表明存在可迁移的实现层面增益。
英文摘要:
Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirements, each often tied to localized code regions such as event handlers, state updates, DOM fragments, or CSS selectors. Standard GRPO collapses these structured outcomes into a single sequence-level reward and applies the resulting advantage uniformly to all tokens, weakening credit assignment. We propose \textbf{Rubric-to-Code Credit Assignment} (RCCA), a reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code. RCCA builds training tasks around explicit functional rubrics, uses a hierarchical reward to separate format, source-code, runtime, and functional failures, and aligns evaluator-generated textual attributions with responsible code spans and generated tokens. The resulting model, \textbf{Ling-RCCA-Flash}, scores 41.25 on MiniAppBench, improving Ling-3.0-Flash by 32.20 points and slightly surpassing Claude Opus 4.5. It also reaches 76.19 on ArtifactsBench, improving the SFT model by 4.48 points and establishing a new top score under the official ArtifactsBench leaderboard setting by surpassing the GPT-5 score by 3.64 points, suggesting transferable implementation-level gains.