arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30968cs.CLcs.AI

CogEvol:面向高效且可靠的学习环境生成

CogEvol: Towards Efficient and Reliable Learning Environment Generation

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Binglin Liu, Ye He, Danqi Zhe… 展开作者

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出专门用于学习环境生成的CogEvol模型,可将课程简介一次性转化为学习产物,在22万条生产请求中表现高效可靠,参数规模远小于旗舰编码模型,还可在国产昇腾加速器运行,降低AI原生教育单位成本。

中文摘要 AI 辅助

我们提出了CogEvol,这是一类专门针对学习环境生成任务训练的模型:能够一次性将课程简介转化为成品学习产物(结构化JSON幻灯片或自包含的交互式HTML页面)。在22万条生产请求中,CogEvol生成幻灯片的中位数耗时为17秒,生成交互式页面的中位数耗时为59秒,替代了原本需要数分钟的多轮智能体支架式流程。可靠性是通过强制手段实现的,而非依赖运气:一条基于生产环境的数据流将真实故障转化为53687个验证过的监督微调(SFT)样本,混合规则加视觉语言模型(VLM)的奖励机制驱动基于组相对策略优化(GRPO)的强化学习,我们在发现并修复了一个生成视觉上可信但无法运行的游戏的奖励作弊事件后,进一步强化了该机制。CogEvol-27B在幻灯片质量指标上得分为83.7,在包含500个案例的交互式HTML基准测试中得分为63.7,其参数规模比旗舰级编码模型少26.9倍,且与OpenMAIC团队合作,为其实时生产流量提供服务。CogEvol-4B以Apache 2.0许可证公开发布于指定网址,外部旗舰模型在相同测试套件和统一测试框架下接受评估。支架编辑将交互式页面生成的成本进一步降低了约76%,整套模型栈可在国产昇腾(Ascend)加速器上运行,性能与A800 GPU达到应用级 parity(同等水平),降低了规模化AI原生教育的单位成本。

英文摘要

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.

发表机构

  • CogEvol Inc.(CogEvol公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑