Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
Nested-ReFT:通过离策略展开实现大语言模型微调的高效强化学习
Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi
机构
*
Mila - Québec AI Institute, Canada(魁北克人工智能研究所)
;
Huawei Noah's Ark Lab (Montreal Research Center), Canada(华为诺亚实验室(蒙特利尔研究中心))
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
专题命中
效率与部署
:large language model(title);language model(title);post-training(abstract);分类 cs.CL、cs.AI、cs.LG
Memorization Dynamics in Knowledge Distillation for Language Models
知识蒸馏中语言模型的知识记忆动态
Jaydeep Borkar, Karan Chadha, Niloofar Mireshghallah, Yuchen Zhang, Irina-Elena Veliche, Archi Mitra, David A. Smith, Zheng Xu, Diego Garcia-Olano
机构
*
Meta Superintelligence Labs(Meta超智能实验室)
;
Meta Central Applied Science(Meta中央应用科学)
;
FAIR at Meta(Meta的FAIR)
;
Carnegie Mellon University(卡内基梅隆大学)
专题命中
效率与部署
:language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.CL
LLM-Extracted Covariates for Clinical Causal Inference: Rethinking Integration Strategies
基于大语言模型的协变量用于临床因果推断:重新思考整合策略
Lei Liu, Jialin Chen, Kathy Macropol
机构
*
Department of Computer Science and Mathematics(计算机科学与数学系)
;
Arcadia University(阿卡迪亚大学)
;
Department of Computer Science(计算机科学系)
;
Yale University(耶鲁大学)
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
从一开始就注定失败:通过召回控制的探测级联实现大语言模型智能体情节的早期终止
Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Jinghao Lin, Xuan Wang, Hao Sun
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Duke University(杜克大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
College of Computer Science, Zhejiang University(浙江大学计算机学院)
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI