arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-10-09 至 2025-10-09 共收录 8 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 8 篇

2510.06652 2025-10-09 cs.CL 91%

Aligning Large Language Models via Fully Self-Synthetic Data

Shangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen, Yu Meng

机构 * University of California, Riverside(加州大学河滨分校) University of Virginia(弗吉尼亚大学)

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06988 2025-10-09 cs.CV 78%

No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts

Girolamo Macaluso, Lorenzo Mandelli, Mirko Bicchierai, Stefano Berretti, Andrew D. Bagdanov

机构 * University of Florence(佛罗伦萨大学)

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22296 2025-10-09 cs.CL cs.LG 76%

360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training

Haosheng Zou, Xiaowei Lv, Shousheng Jia, Lin Li, Xiaochun Gong, Xiangzheng Zhang

机构 * Qiyuan Tech(启元科技) Renmin University(中国人民大学)

专题命中 后训练与偏好优化 :post-training(title);分类 cs.CL、cs.LG

Comments v2: sp for vlms; code at https://github.com/Qihoo360/360-LLaMA-Factory

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06391 2025-10-09 cs.CL cs.AI 73%

Reward Model Perspectives: Whose Opinions Do Reward Models Reward?

Elle

机构 * Elle University of Oxford(埃勒 奥克舍大学)

专题命中 后训练与偏好优化 :language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

Comments Published at EMNLP 2025 under the full author name "Elle"

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06878 2025-10-09 cs.AI 70%

TGPR: Tree-Guided Policy Refinement for Robust Self-Debugging of LLMs

Daria Ozerova, Ekaterina Trofimova

机构 * HSE University(俄罗斯莫斯科国立高等经济学院)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06866 2025-10-09 cs.CL 70%

Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding

Wafaa Mohammed, Vlad Niculae, Chrysoula Zerva

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03905 2025-10-09 cs.CL 70%

Sotopia-RL: Reward Design for Social Intelligence

Haofei Yu, Zhengyang Qi, Yining Zhao, Kolby Nottingham, Keyang Xuan, Bodhisattwa Prasad Majumder, Hao Zhu, Paul Pu Liang, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Irvine(加州大学尔湾分校) Allen Institute for Artificial Intelligence(人工智能艾伦研究所) Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06629 2025-10-09 cs.CR cs.CV cs.LG 57%

Unsupervised Backdoor Detection and Mitigation for Spiking Neural Networks

Jiachen Li, Bang Wu, Xiaoyu Xia, Xiaoning Liu, Xun Yi, Xiuzhen Zhang

机构 * RMIT University(皇家墨尔本理工大学)

专题命中 后训练与偏好优化 :post-training(abstract);分类 cs.LG

Comments To appear in The 28th International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏