SPADE: Self-Play in Adaptive Synthetic Executable Environments
SPADE:自适应合成可执行环境中的自博弈
Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
机构
*
University of Washington(华盛顿大学)
;
Stanford University(斯坦福大学)
;
Northeastern University(东北大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
National University of Singapore(新加坡国立大学)
;
Seoul National University(首尔大学)
;
Stevens Institute of Technology(史蒂文斯理工学院)
;
University of Chicago(芝加哥大学)
Rethinking Reverse KL as Adaptive Entropy Distillation
将反向KL重新思考为自适应熵蒸馏
Shizhen Li, Zhiyu Shen, Yuyin Lu, Yunhe Pang, Jielin Song, Yanghui Rao, Fu Lee Wang
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院)
Enhancing LLM Metacognition via Cognitive Pairwise Training
通过认知成对训练增强LLM元认知
Weitao Li, Hao Zhou, Xuanyu Lei, Fandong Meng, Yuanhang Liu, Jingyi Ren, Ante Wang, Xiaolong Wang, Yuanchi Zhang, Fuwen Luo, Guangwen Yang, Lin Gan, Weizhi Ma, Yang Liu
机构
*
National Engineering Laboratory for Intelligent Information Processing, Academy of Mathematics and Physics, Chinese Academy of Sciences(智能信息处理国家工程实验室,中国科学院数学物理研究所)
;
University of Science and Technology of China(中国科学技术大学)
CommentsWithdrawn due to serious concerns regarding the authenticity and accuracy of the listed authorship. The identity of one or more listed authors cannot presently be verified, and the author list may not represent distinct contributors. The manuscript is withdrawn pending institutional review
CommentsThis is an extended and corrected version of the paper presented at the 22nd International Conference on Principles of Knowledge Representation and Reasoning (KR 2025); see the appendix for details
MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
MAVEN-T:用于实时多智能体轨迹预测的强化异构蒸馏
Wenchang Duan, Zhenguo Gao, Jinguo Xian, Yi Shi
机构
*
School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院)
;
Bio-X Institutes, Key Laboratory for the Genetics of Developmental and Neuropsychiatric Disorders, Shanghai Jiao Tong University(上海交通大学Bio-X研究院、发育与神经精神疾病遗传学重点实验室)
;
Shanghai Key Laboratory of Psychotic Disorders, Brain Science and Technology Research Center, Shanghai Jiao Tong University(上海精神疾病重点实验室、脑科学与技术研究中心,上海交通大学)
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
HaReCAP:面向递归大语言模型智能体的习惯性动作 grounding 方法
Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang
机构
*
North China Institute of Computer System Engineering(华北计算机系统工程研究所)
;
University of Science and Technology of China(中国科学技术大学)
;
China Information Security Research Institute Co., Ltd.(中国信息安全研究院有限公司)
CommentsWithdrawn due to serious concerns regarding the authenticity and accuracy of the listed authorship. The identity of one or more listed authors cannot presently be verified, and the author list may not represent distinct contributors. The manuscript is withdrawn pending institutional review