DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.LG
Comments This paper is subsumed by a later paper of ours: arXiv:2411.00361
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.LG
Comments This paper is subsumed by a later paper of ours: arXiv:2411.00361
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI
Comments Accepted by AAAI 2025
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.LG
Comments Mechanistic Interpretability workshop at ICML 2024; Main conference poster at NeurIPS 2024
专题命中 后训练与偏好优化 :LLM(title,abstract);分类 cs.LG
Comments accepted by NeurIPS 2024
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL
Comments EMNLP 2024
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.LG
Comments NeurIPS accepted version
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.LG
Comments Accepted to NeurIPS 2024
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
Comments 9 pages, 14 images, 6 tables
Journal ref Phys Rev A 110 042409 (2024)
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.LG
Comments EMNLP 2024 Main
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
Comments Accepted by ECCV 2024
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
Comments Accepted by ECCV2024
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.LG
Comments Shorter version accepted to ICML 2024
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.AI
Comments Accepted by CVPR2024
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL
Comments The project is available at https://github.com/NJUNLP/MAPO
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL
Comments Accepted by COLING 2024
专题命中 后训练与偏好优化 :RLHF(title);preference optimization(abstract);分类 cs.LG
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.LG
Comments 14 pages, 3 figures
Journal ref ICML2023 Interactive Learning from Implicit Human Feedback Workshop
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.AI
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.LG
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.CL
Comments Accepted by ACL 2023 Findings, Long Paper
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.AI
Comments 15 pages