机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
Sun Yat-sen University(中山大学)
;
Central South University(中南大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Tencent Inc.(腾讯公司)
State2State: Environment-Derived Mid-Training for LLM Agents
State2State:面向大语言模型智能体的环境衍生式中间训练
Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu, Peng Li, Ming Yan, Jieping Ye, Ya-Qin Zhang, Yang Liu
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Institute for AI, Tsinghua University(清华大学人工智能研究院)
;
Institute of Intelligent Computing, Alibaba Group(阿里巴巴集团智能计算研究院)
CommentsDisclaimer. This manuscript is provided as an arXiv preprint to establish a public record of the NeuroSynth continual reinforcement learning architecture and its evaluation on the NeuroMaze-CL benchmark. This full manuscript has been submitted to the Journal of High School Science for peer review