arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-10-29 至 2025-10-29 共收录 5 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 5 篇

2510.24700 2025-10-29 cs.LG cs.AI cs.IT math.IT stat.ML 88%

Greedy Sampling Is Provably Efficient for RLHF

Di Wu, Chengshuai Shi, Jing Yang, Cong Shen

机构 * Electrical and Computer Engineering University of Virginia(电气与计算机工程大学弗吉尼亚大学) Princeton Language and Intelligence Princeton University(普林斯顿语言与智能普林斯顿大学)

专题命中 后训练与偏好优化 :RLHF(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24695 2025-10-29 cs.CL 87%

AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis

Xuanzhong Chen, Zile Qiao, Guoxin Chen, Liangcai Su, Zhen Zhang, Xinyu Wang, Pengjun Xie, Fei Huang, Jingren Zhou, Yong Jiang

机构 * Tongyi Lab(通义实验室) Alibaba Group(阿里巴巴集团)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

Comments https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01381 2025-10-29 cs.CL 83%

AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation

Yilong Lai, Jialong Wu, Zhenglin Wang, Deyu Zhou

机构 * School of Computer Science and Engineering, Key Laboratory of Computer Network and Information Integration, Ministry of Education, Southeast University(计算机科学与工程学院、计算机网络与信息集成重点实验室、教育部、东南大学)

专题命中 后训练与偏好优化 :prompting(title,abstract);LLM(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23751 2025-10-29 cs.LG cs.AI stat.ML 73%

Debiasing Reward Models by Representation Learning with Guarantees

Ignavier Ng, Patrick Blöbaum, Siddharth Bhandari, Kun Zhang, Shiva Kasiviswanathan

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon(亚马逊) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德智能大学)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22820 2025-10-29 cs.LG cs.AI econ.TH stat.ML 62%

Preference Learning with Response Time: Robust Losses and Guarantees

Ayush Sawarni, Sahasrajit Sarmasarkar, Vasilis Syrgkanis

专题命中 后训练与偏好优化 :foundation model(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏