Reinforced Language Models for Sequential Decision Making
专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);large language model(abstract);post-training(abstract)
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);large language model(abstract);post-training(abstract)
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) ; CASIA(中国科学院自动化所) ; Alibaba Cloud(阿里云) ; School of Intelligence Science and Technology(智能科学与技术学院) ; CAIR(中国科学院香港创新研究院)
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.AI
机构 * 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited ; 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles ; 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute ; 9Noah's Ark Lab 10CG Matrix Technology Limited
专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);preference optimization(abstract);分类 cs.AI、cs.LG
Comments Accepted at the 34th ACM International Conference on Information and Knowledge Management (CIKM2025)
专题命中 后训练与偏好优化 :preference optimization(title,abstract)
Comments Project Page: https://haroldchen19.github.io/PhysHPO-Page/