arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-20 至 2025-11-20 共收录 4 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4 篇

2511.15392 2025-11-20 cs.CL cs.AI 91%

DEPO: Dual-Efficiency Preference Optimization for LLM Agents

DEPO:用于LLM代理的双效偏好优化

Sirui Chen, Mengshi Zhao, Lei Xu, Yuying Zhao, Beier Zhu, Hanwang Zhang, Shengjie Zhao, Chaochao Lu

专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

AI总结 DEPO通过双效偏好优化方法提升LLM代理的效率和性能,减少token和步骤消耗,同时在多个基准测试中表现优异。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15256 2025-11-20 cs.LG cs.CV 77%

GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning

Yanchen Xu, Ziheng Jiao, Hongyuan Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院(iOPEN),西北工业大学) HuaWei Technologies Co., Ltd.(华为技术有限公司) The University of Hong Kong(香港大学)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20964 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 62%

Core Safety Values for Provably Corrigible Agents

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 后训练与偏好优化 :RLHF(abstract);分类 cs.AI、cs.LG

Comments 14 pages. To appear in AAAI 2026 Machine Ethics Workshop (W37) Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15038 2025-11-20 cs.SD cs.AI eess.AS 57%

Aligning Generative Music AI with Human Preferences: Methods and Challenges

Dorien Herremans, Abhinaba Roy

专题命中 后训练与偏好优化 :preference optimization(abstract);分类 cs.AI

Comments Accepted at the AAAI-2026 Senior Member Track

详情

展开后加载摘要…

URL PDF HTML 收藏