RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
RE-PO:一种用于大语言模型对齐的鲁棒增强策略优化通用框架
Xiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long, Michiel A. Bakker, Yu Wang, Chao Yu
机构
*
IDSS, Massachusetts Institute of Technology(IDSS,麻省理工学院)
;
EE, Tsinghua University(电子工程系,清华大学)
;
Li Auto Inc.(力汽车公司)
;
SIGS, Tsinghua University(系统工程系,清华大学)
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
通过元权重在线采样对齐:弥合数据生成与偏好优化之间的差距
Junming Yang, Ning Xu, Biao Liu, Shiqi Qiao, Xin Geng
机构
*
School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用国家重点实验室)
机构
*
Qwen Large Model Application Team, Alibaba(阿里巴巴文勤大模型应用团队)
;
Beijing University Of Posts and Telecommunications(北京邮电大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
Wenzhe Zhao, Yang Zhao, Ganchao Liu, Zhiyu Jiang, Dandan Ma, Zihao Li, Xuelong Li
机构
*
School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China(人工智能学院、光学与电子学(iOPEN)、西北工业大学,西安 710072,中国)
;
Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究所(TeleAI)、中国电信,中国)
TSC: Topology-Conditioned Stackelberg Coordination for Multi-Agent Reinforcement Learning in Interactive Driving
TSC:基于拓扑条件的Stackelberg协调用于交互驾驶中的多智能体强化学习
Xiaotong Zhang, Gang Xiong, Yuanjing Wang, Siyu Teng, Alois Knoll, Long Chen
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Department of Natural Sciences, University of Durham(达勒姆大学自然科学系)
;
College of Civil and Transportation Engineering, Shenzhen University(深圳大学土木与交通工程学院)
;
Chair of Robotics, Artificial Intelligence and Realtime Systems, Technical University of Munich(慕尼黑技术大学机器人、人工智能与实时系统教授职位)
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(摩埃关键实验室,中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
National University of Singapore(新加坡国立大学)
;
Tianjin University(天津大学)
Resilient Strategies for Stochastic Systems: How Much Does It Take to Break a Winning Strategy?
随机系统中的稳健策略:打破获胜策略需要多大的代价?
Kush Grover, Markel Zubia, Debraj Chakraborty, Muqsit Azeem, Nils Jansen, Jan Kretinsky
机构
*
Ruhr University Bochum Germany
;
Nanyang Technological University, Singapore
;
Technical University of Munich \& University of Konstanz Germany
;
Ruhr University Bochum \& Radboud University Nijmegen Germany
;
Masaryk University Czech Republic
;
Ruhr University Bochum
;
Technical University of Munich \& University of Konstanz
;
Ruhr University Bochum \& Radboud University Nijmegen
;
Masaryk University
CommentsTo appear in Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), Paphos, Cyprus, May 25-29, 2026
机构
*
Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室)
;
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences(脑认知实验室,中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Long-term AI(长期人工智能)
Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
统一的生成与理解模型能否在不同输出模态间保持语义等价性?
Hongbo Jiang, Jie Li, Yunhang Shen, Pingyang Dai, Xing Sun, Haoyu Cao, Liujuan Cao
机构
*
Tencent Youtu Lab(腾讯云图实验室)
;
Xiamen University(厦门大学)
;
Computing Lab, Department of Artificial Intelligence, School of Informatics(计算实验室,人工智能系,信息学院)