机构
*
Technische Universität Berlin(柏林技术大学)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
University of Duisburg-Essen(杜伊斯堡- Essen大学)
;
LMU Munich(慕尼黑大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Saarland Informatics Campus(萨尔兰州信息学校区)
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究院)
;
Centre for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
专题命中
后训练与偏好优化
:post-training(title,abstract);large language model(abstract_cn);language model(abstract_cn);分类 cs.AI
GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints
GRIP:通过几何路由约束实现混合专家的算法无关机器反学习
Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li
机构
*
School of Computer Science(计算机科学学院)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Department of Electrical Engineering(电气工程系)
;
Tsinghua University(清华大学)
;
School of Electrical and Computer Engineering(电气与计算机工程学院)
专题命中
后训练与偏好优化
:large language model(abstract);language model(abstract);post-training(abstract);分类 cs.AI、cs.LG
AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
你的强化学习奖励函数是你的最佳搜索PRM:统一强化学习与基于搜索的文本生成
Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han, Zhang-Wei Hong, Tong Che, Dimitris N. Metaxas
机构
*
Rutgers University(新泽西州立大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Connecticut(康涅狄格大学)
;
Fudan University(复旦大学)
;
NVIDIA Research(NVIDIA研究)
;
Red Hat AI Innovation(红帽AI创新)
;
MIT-IBM Watson AI Lab(MIT-IBM沃森AI实验室)
;
Massachusetts Institute of Technology(麻省理工学院)
专题命中
后训练与偏好优化
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
智能体游戏开发:用于扩展世界模型的可验证轨迹数据引擎
Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, Yixing Ma, Zhenglin Wan, Kaipeng Zhang, Wangbo Zhao, Yang You
机构
*
Cardinal AI Lab(卡迪纳尔人工智能实验室)
;
University of California, Berkeley(加州大学伯克利分校)
;
Hong Kong University of Science and Technology(香港科技大学)
;
National University of Singapore(新加坡国立大学)
;
HPC-AI Lab(高性能计算与人工智能实验室)
机构
*
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进创新中心,人工智能学院,北京航空航天大学)
;
Tsinghua University(清华大学)