arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4541 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4541 篇

2412.05095 2025-10-21 cs.CV 78%

SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization

Xiaofeng Tan, Hongsong Wang, Xin Geng, Pan Zhou

机构 * Department of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程系) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China(新一代人工智能技术及其交叉应用国家重点实验室(东南大学))

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Journal ref Advances in Neural Information Processing Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09056 2025-10-13 cs.CV 78%

Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion

Junhyeok Lee, Hyunwoong Kim, Hyungjin Chung, Heeseong Eom, Joon Jang, Chul-Ho Sohn, Kyu Sung Choi

机构 * College of Medicine, Seoul National University, Seoul, Republic of Korea(首尔国立大学医学院) Department of Radiology, Seoul National University Hospital(首尔国立大学医院放射科) Department of Biomedical Sciences, Seoul National University(首尔国立大学生物医学科学系)

专题命中 后训练与偏好优化 :post-training(title,abstract)

Comments MICCAI 2025, Lecture Notes in Computer Science Vol. 15961

Journal ref Med Image Comput Comput Assist Interv. LNCS 15961, 282-291, Springer, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06988 2025-10-09 cs.CV 78%

No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts

Girolamo Macaluso, Lorenzo Mandelli, Mirko Bicchierai, Stefano Berretti, Andrew D. Bagdanov

机构 * University of Florence(佛罗伦萨大学)

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25718 2025-10-01 cs.RO 78%

VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning

Si-Cheng Wang, Tian-Yu Xiang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Ao-Qun Jin, Zeng-Guang Hou

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23958 2025-09-30 cs.CV 78%

Reinforcement Learning with Inverse Rewards for World Model Post-training

Yang Ye, Tianyu He, Shuo Yang, Jiang Bian

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02588 2025-09-29 cs.CV 78%

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

Kyungmin Lee, Xiaohang Li, Qifei Wang, Junfeng He, Junjie Ke, Ming-Hsuan Yang, Irfan Essa, Jinwoo Shin, Feng Yang, Yinxiao Li

机构 * Google DeepMind(谷歌DeepMind) KAIST(韩国科学技术院) Google(谷歌) Google Research(谷歌研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments CVPR 2025, Project page: https://kyungmnlee.github.io/capo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18928 2025-09-24 eess.AS 78%

Direct Preference Optimization for Speech Autoregressive Diffusion Models

Zhijun Liu, Dongya Jia, Xiaoqiang Wang, Chenpeng Du, Shuai Wang, Zhuo Chen, Haizhou Li

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18463 2025-09-10 cs.CV 78%

DIP: Unsupervised Dense In-Context Post-training of Visual Representations

Sophia Sirko-Galouchenko, Spyros Gidaris, Antonin Vobecky, Andrei Bursuc, Nicolas Thome

机构 * Sorbonne Université, CNRS, ISIR(索邦大学、国家科学研究中心、ISIR) FEE CTU(捷克技术大学电子工程系) CIIRC CTU Prague(捷克技术大学普拉格研究所) Institut universitaire de France (IUF)(法国国家科学院)

专题命中 后训练与偏好优化 :post-training(title,abstract)

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11255 2025-08-18 cs.CV 78%

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

MengChao Wang, Qiang Wang, Fan Jiang, Mu Xu

机构 * Project Leader(项目负责人)

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments https://fantasy-amap.github.io/fantasy-talking2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10858 2025-08-15 cs.CV 78%

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang, Ser-Nam Lim

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Project Page: https://haroldchen19.github.io/PhysHPO-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18433 2025-07-25 eess.IV cs.CV 78%

DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis

Minxi Ouyang, Lianghui Zhu, Yaqing Bao, Qiang Huang, Jingli Ouyang, Tian Guan, Xitong Ling, Jiawen Li, Song Duan, Wenbin Dai, Li Zheng, Xuemei Zhang, Yonghong He

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Department of Pathology, Liuzhou People’s Hospital Affiliated to Guangxi Medical University(广西医科大学柳州市人民医院病理科) Department of Immunology, College of Basic Medical Sciences, China Medical University(中国医科大学基础医学学院免疫科) Greater Bay Area Center for Medical Device Evaluation and Inspection.NMPA(粤港澳大湾区医疗器械评价和检验中心.NMPA) Shenzhen Shengqiang Technology Co., Ltd.(深圳盛强科技有限公司) Department of Pathology, Chongqing University Affiliated Three Gorges Hospital(重庆大学附属第三人民医院病理科)

专题命中 后训练与偏好优化 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18655 2025-06-24 cs.CV 78%

RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

Wenxu Qian, Chaoyue Wang, Hou Peng, Zhiyu Tan, Hao Li, Anxiang Zeng

机构 * Fudan University(复旦大学) Shopee Inc(Shopee公司)

专题命中 后训练与偏好优化 :preference optimization(title);post-training(abstract)

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07188 2025-06-10 cs.CV 78%

Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks

Ni Ding, Lei He, Shengbo Eben Li, Keqiang Li

机构 * School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China(清华大学车辆与移动系统学院) State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China(清华大学智能绿色车辆与移动系统国家重点实验室) School of Mathematics and Statistics, Beijing Institute of Technology, Beijing 100081, China(北京理工大学数学与统计学院)

专题命中 后训练与偏好优化 :post-training(title,abstract)

Comments 13 pages, 7 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02698 2025-06-09 cs.CV 78%

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

Yunhong Lu, Qichao Wang, Hengyuan Cao, Xiaoyin Xu, Min Zhang

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11475 2025-06-04 cs.SE 78%

Focused-DPO: Enhancing Code Generation Through Focused Preference Optimization on Error-Prone Points

Kechi Zhang, Ge Li, Jia Li, Yihong Dong, Jia Li, Zhi Jin

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Camera Ready version for ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22002 2025-05-29 cs.CV 78%

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

Zijing Hu, Fengda Zhang, Kun Kuang

机构 * Zijing Hu(Hu Zijing) Fengda Zhang(Zhang Fengda) Kun Kuang(Kuang Kun)

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18503 2025-05-27 cs.CV 78%

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

Aofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Fenglong Ma

专题命中 后训练与偏好优化 :language model(title,abstract)

Comments Accepted to ACL2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10250 2025-05-23 cs.CV 78%

ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization

Wenhao Shen, Wanqi Yin, Xiaofeng Yang, Cheng Chen, Chaoyue Song, Zhongang Cai, Lei Yang, Hao Wang, Guosheng Lin

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Accepted to ICML 2025. Code: https://github.com/shenwenhao01/ADHMR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11245 2025-05-19 cs.CV 78%

Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models

Fu-Yun Wang, Yunhao Shui, Jingtan Piao, Keqiang Sun, Hongsheng Li

机构 * MMLab, CUHK, Hong Kong(CUHK的MMLab, 香港) Shanghai Jiang Tong University, Shanghai(上海江 Tong大学, 上海) CPII under InnoHK, Hong Kong(InnoHK下的CPII, 香港)

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13548 2025-05-13 cs.CL cs.AI cs.LG stat.ML 78%

Towards Understanding Sycophancy in Language Models

Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, Ethan Perez

专题命中 后训练与偏好优化 :language model(title);分类 cs.CL、cs.AI、cs.LG

Comments 32 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12900 2025-04-18 cs.MM cs.IR 78%

FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization

Mingzhe Yu, Yunshan Ma, Lei Wu, Changshuo Wang, Xue Li, Lei Meng

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Accepted by SIGIR'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08542 2025-04-14 cs.CV 78%

Discriminator-Free Direct Preference Optimization for Video Diffusion

Haoran Cheng, Qide Dong, Liang Peng, Zhizhou Sha, Weiguo Feng, Jinghui Xie, Zhao Song, Shilei Wen, Xiaofei He, Boxi Wu

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments arXiv admin note: text overlap with arXiv:2412.14167 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08977 2025-03-26 cs.CV 78%

Text-driven 3D Human Generation via Contrastive Preference Optimization

Pengfei Zhou, Xukun Shen, Yong Hu

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments 10+2

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09595 2025-03-13 cs.CV 78%

PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

Chenyu Li, Oscar Michel, Xichen Pan, Sainan Liu, Mike Roberts, Saining Xie

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04628 2025-02-10 cs.CV 78%

AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers

Runqing Jiang, Ye Zhang, Longguang Wang, Pengpeng Yu, Yulan Guo

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01690 2025-02-05 cs.CV 78%

HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment

Lifan Jiang, Boxi Wu, Jiahui Zhang, Xiaotong Guan, Shuang Chen

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09859 2025-01-30 cs.ET 78%

Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-Training

Corey Lammie, Athanasios Vasilopoulos, Julian Büchel, Giacomo Camposampiero, Manuel Le Gallo, Malte Rasch, Abu Sebastian

专题命中 后训练与偏好优化 :post-training(title,abstract)

Comments Accepted at 2024 IEEE International Symposium on Circuits and Systems (ISCAS)

Journal ref 2024 IEEE International Symposium on Circuits and Systems (ISCAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09846 2025-01-23 cs.SE 78%

MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning

Jinhan Kim, Nargiz Humbatova, Gunel Jahangirova, Shin Yoo, Paolo Tonella

专题命中 后训练与偏好优化 :post-training(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18013 2024-10-31 cs.CV 78%

Scalable Ranked Preference Optimization for Text-to-Image Generation

Shyamgopal Karthik, Huseyin Coskun, Zeynep Akata, Sergey Tulyakov, Jian Ren, Anil Kag

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

Comments Project Page: https://snap-research.github.io/RankDPO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06737 2024-10-25 cs.IR 78%

Post-Training Attribute Unlearning in Recommender Systems

Chaochao Chen, Yizhao Zhang, Yuyuan Li, Jun Wang, Lianyong Qi, Xiaolong Xu, Xiaolin Zheng, Jianwei Yin

专题命中 后训练与偏好优化 :post-training(title,abstract)

Comments Accepted by TOIS

详情

展开后加载摘要…

URL PDF HTML 收藏