arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-11 至 2025-11-11 共收录 9 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 9 篇

2511.07070 2025-11-11 cs.AI cs.LG 91%

RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services

Fei Zhao, Chonggang Lu, Haofu Qian, Fangcheng Shi, Zijie Meng, Jianzhao Huang, Xu Tang, Zheyong Xie, Zheyu Ye, Zhe Xu, Yao Hu, Shaosheng Cao

机构 * NLP Team, Xiaohongshu Inc.(小红书研究院自然语言处理团队)

专题命中 后训练与偏好优化 :LLM(title,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06023 2025-11-11 cs.CL 89%

Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data

Deng Yixuan, Ji Xiaoqiang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01450 2025-11-11 cs.CV cs.AI 88%

Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation

Jie Du, Xinyu Gong, Qingshan Tan, Wen Li, Yangming Cheng, Weitao Wang, Chenlu Zhan, Suhui Wu, Hao Zhang, Jun Zhang

专题命中 后训练与偏好优化 :SFT(title,abstract);preference optimization(title,abstract);分类 cs.AI

Comments The paper is withdrawn due to the need for further revision and verification of experimental results. A revised version will be resubmitted once the updates are completed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06722 2025-11-11 cs.CV cs.AI cs.CL 88%

Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View

Jianyu Qi, Ding Zou, Wenrui Yan, Rui Ma, Jiaxu Li, Zhijie Zheng, Zhiguo Yang, Rongchang Zhao

专题命中 后训练与偏好优化 :post-training(title,abstract);large language model(abstract);language model(abstract);SFT(abstract)

Comments Accpeted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09039 2025-11-11 cs.CL 85%

Atomic Consistency Preference Optimization for Long-Form Question Answering

Jingfeng Chen, Raghuveer Thirukovalluru, Junlin Wang, Kaiwei Luo, Bhuwan Dhingra

专题命中 后训练与偏好优化 :preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments 13 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07384 2025-11-11 cs.CL cs.AI cs.LG 85%

Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence

Sean McLeish, Ang Li, John Kirchenbauer, Dayal Singh Kalra, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Jonas Geiping, Tom Goldstein, Micah Goldblum

机构 * University of Maryland(马里兰大学) New York University(纽约大学) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) University of North Carolina(北卡罗来纳大学) ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center(图宾根ELLIS研究所、马克斯·普朗克智能系统研究所、图宾根人工智能中心) Columbia University(哥伦比亚大学)

专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG

Comments code: https://github.com/mcleish7/retrofitting-recurrence, models: https://huggingface.co/collections/tomg-group-umd/retrofitting-recurrence

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06618 2025-11-11 cs.AI cs.CL cs.LG cs.SE 80%

GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization

Moriya Dechtiar, Daniel Martin Katz, Mari Sundaresan, Sylvain Jaume, Hongming Wang

机构 * Harvard University(哈佛大学) Illinois Tech - Chicago Kent College of Law(伊利诺伊理工学院-芝加哥肯特法学院) CLTDS, Bucerius Law School(CLTDS,布塞里乌斯法学院) Yong Pung How School of Law, Singapore Management University(永丰何法学院,新加坡管理大学) CodeX - The Stanford Center for Legal Informatics, Stanford University(CodeX-斯坦福法律信息中心,斯坦福大学) Georgetown University(乔治城大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05616 2025-11-11 cs.CV cs.AI 79%

Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization

Connor Dunlop, Matthew Zheng, Kavana Venkatesh, Pinar Yanardag

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI

Comments Published at NeurIPS'25 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03262 2025-11-11 cs.CL cs.LG 79%

REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Jian Hu, Jason Klein Liu, Haotian Xu, Wei Shen

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);RLHF(abstract);分类 cs.CL、cs.LG

Comments refactor

详情

展开后加载摘要…

URL PDF HTML 收藏