DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
DreamPRM-Code: 一种基于函数作为步骤的进程奖励模型与标签校正用于LLM编码
机构 * University of California, San Diego(加州大学圣迭戈分校)
专题命中 代码生成 :code generation(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 DreamPRM-Code通过链式函数提示策略和元学习校正机制,提升LLM在编码任务中的表现,达到LiveCodeBench的最高准确率