Comments21 pages, 6 figures, 9 tables. Reports a 78-day deployment across three heterogeneous industrial recommender business lines (1,624 CLI-tool dispatches). Companion paper: AutoResearch (P3b), which instantiates the same substrate for autonomous research
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
DiDPO:用于编码智能体训练的差异内差异策略优化
Xucong Wang, Zhe Zhao, Liheng Yu, Di Wu, Xiaofeng Cao, Pengkun Wang
机构
*
University of Science and Technology of China (USTC)(中国科学技术大学)
;
Stanford University(斯坦福大学)
;
Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院)
;
Tongji University(同济大学)
EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
EvolveNet:智能体自我改进的协作式工具链进化
Jun Nie, Yonggang Zhang, Qianshu Cai, Yiu-ming Cheung, Xinmei Tian, Bo Han
机构
*
Hong Kong Baptist University(香港浸会大学)
;
University of Science and Technology of China(中国科学技术大学)
;
The Hong Kong University of Science and Technology(香港科技大学)