arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-04 至 2025-12-04 共收录 166 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 16 篇

2512.03442 2025-12-04 cs.CL 85%

PretrainZero: Reinforcement Active Pretraining

PretrainZero:强化主动预训练

Xingrun Xing, Zhiyuan Fan, Jie Lou, Guoqi Li, Jiajun Zhang, Debing Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Xiaohongshu Inc.(小红书公司)

专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract);post-training(abstract);分类 cs.CL

AI总结 PretrainZero通过主动预训练和自监督学习方法,提升通用推理能力,并在多个基准测试中取得显著成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03307 2025-12-04 cs.LG cs.AI 85%

Robust Tabular Foundation Models

鲁棒表格基础模型

Matthew Peroni, Franck Le, Vadim Sheinin

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本文提出鲁棒表格基础模型(RTFM),通过参数化生成器分布实现对抗鲁棒性,提升表格基础模型的基准性能。

Comments Shaping Responsible Synthetic Data in the Era of Foundation Models, AAAI 2026

Journal ref Shaping Responsible Synthetic Data in the Era of Foundation Models, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03445 2025-12-04 cs.CV cs.AI 83%

Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation

多方面知识增强的医学视觉-语言预训练与多代理数据生成

Xieji Li, Siyuan Yan, Yingsheng Liu, H. Peter Soyer, Monika Janda, Victoria Mar, Zongyuan Ge

机构 * Department of Data Science and AI, Faculty of Information Technology, Monash University(数据科学与人工智能系,信息科技学院,墨尔本大学) Victorian Melanoma Service, Alfred Health(维多利亚黑色素瘤服务,阿尔弗雷德健康) Frazer Institute, The University of Queensland, Dermatology Research Centre(弗雷泽研究所,昆士兰大学,皮肤科研究中心)

专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract);分类 cs.AI

AI总结 本研究提出一种多代理数据生成与多方面知识增强的医学视觉-语言预训练框架,通过提升数据质量和细粒度对齐,实现零样本性能的突破。

Comments 10 pages. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04005 2025-12-04 cs.CV cs.LG cs.RO 83%

LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving

LargeAD: 大规模跨传感器数据预训练用于自动驾驶

Lingdong Kong, Xiang Xu, Youquan Liu, Jun Cen, Runnan Chen, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract);分类 cs.LG

AI总结 LargeAD通过跨传感器数据预训练提升自动驾驶中的三维场景理解,结合多模态对比学习和时间一致性,实现更鲁棒的感知性能。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05127 2025-12-04 eess.IV cs.CV q-bio.QM 82%

PixCell: A generative foundation model for digital histopathology images

PixCell:数字病理图像的生成基础模型

Srikar Yellapragada, Alexandros Graikos, Zilinghan Li, Kostas Triaridis, Varun Belagali, Tarak Nath Nandi, Karen Bai, Beatrice S. Knudsen, Tahsin Kurc, Rajarsi R. Gupta, Prateek Prasanna, Ravi K Madduri, Joel Saltz, Dimitris Samaras

机构 * Stony Brook University(石溪大学) Argonne National Laboratory(阿贡国家实验室) The University of Chicago(芝加哥大学) University of Utah(犹他大学)

专题命中 预训练与数据 :foundation model(title,abstract);language model(abstract)

AI总结 PixCell是首个针对数字病理图像的生成基础模型,通过扩散模型在大规模数据集上训练,实现隐私保护的数据生成和虚拟染色任务,提升病理学研究效率。

Comments Project page - https://histodiffusion.github.io/docs/projects/pixcell

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02512 2025-12-04 cs.CV 78%

Two-Stage Vision Transformer for Image Restoration: Colorization Pretraining + Residual Upsampling

双阶段视觉Transformer用于图像恢复:颜色化预训练+残差上采样

Aditya Chaudhary, Prachet Dev Singh, Ankit Jha

机构 * LNMIIT(拉纳印度理工学院)

专题命中 预训练与数据 :pretraining(title,abstract)

AI总结 本文提出双阶段视觉Transformer方法,通过颜色化预训练和残差上采样提升图像超分辨率性能,实现在DIV2K数据集上的高SSIM和PSNR指标。

Comments Accepted as a Tiny Paper at the 13th Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP 2025), IIT Mandi, India. 3 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03334 2025-12-04 cs.CL 77%

Modeling Topics and Sociolinguistic Variation in Code-Switched Discourse: Insights from Spanish-English and Spanish-Guaraní

在混合语言话语中建模主题与社会语言学变异:来自西班牙-英语和西班牙-瓜拉尼语的洞察

Nemika Tyagi, Nelvin Licona Guevara, Olga Kellert

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究利用大型语言模型构建标注流程,分析西班牙-英语和西班牙-瓜拉尼语混合话语中的主题与社会语言学变异,揭示性别、语言主导地位与话语功能的关系,并扩展了双语现象的定量证据。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03080 2025-12-04 physics.chem-ph cs.AI cs.CE 77%

AtomDisc: An Atom-level Tokenizer that Boosts Molecular LLMs and Reveals Structure--Property Associations

AtomDisc: 一种提升分子大语言模型并揭示结构-性质关联的原子级分词器

Mingxu Zhang, Dazhong Shen, Ying Sun

机构 * The Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能动力部,香港科学与技术大学(广州)) The College of Computer Science and Technology, The Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 AtomDisc通过原子级分词提升分子LLM性能,揭示结构-性质关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16345 2025-12-04 cs.CL 70%

NLP Datasets for Idiom and Figurative Language Tasks

用于隐喻和修辞语言任务的NLP数据集

Blake Matheny, Phuong Minh Nguyen, Minh Le Nguyen, Stephanie Reynolds

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) International College of Technology(国际技术学院)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出用于隐喻和修辞语言任务的NLP数据集,通过构建综合隐喻列表和人工标注数据集,评估预训练语言模型在隐喻识别任务中的表现。

Comments 32 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19153 2025-12-04 cs.LG 70%

Test-Time Training Scaling Laws for Chemical Exploration in Drug Design

药物设计中的测试时间训练扩展规律用于化学探索

Morgan Thomas, Albert Bou, Gianni De Fabritiis

机构 * Computational Science Laboratory, Universitat Pompeu Fabra, Barcelona Biomedical Research Park (PRBB), C Dr. Aiguader 88, 08003 Barcelona, Spain(计算科学实验室,庞培法布拉大学,巴塞罗那生物医学研究公园(PRBB),C Dr. Aiguader 88,08003巴塞罗那,西班牙)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本研究提出通过扩展测试时间训练方法提升化学空间探索效率,引入MolExp基准并验证了扩展TTT的对数线性扩展规律,为生成分子设计提供了可扩展的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12409 2025-12-04 cs.CV 67%

S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing

S5: 远程感知中可扩展的半监督语义分割

Liang Lv, Di Wang, Jing Zhang, Lefei Zhang

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,计算机学院,武汉大学)

专题命中 预训练与数据 :foundation model(abstract);pretraining(abstract)

AI总结 S5提出了一种可扩展的半监督语义分割框架,通过大规模数据集和预训练范式提升遥感任务的性能。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03055 2025-12-04 cs.LG cs.AI 62%

Physics-informed self-supervised learning for predictive modeling of coronary artery digital twins

基于物理的自监督学习用于冠状动脉数字孪生的预测建模

Xiaowu Sun, Thabo Mahendiran, Ortal Senouf, Denise Auberson, Bernard De Bruyne, Stephane Fournier, Olivier Muller, Pascal Frossard, Emmanuel Abbe, Dorina Thanou

机构 * Chair of Mathematical Data Science, EPFL, Lausanne, Switzerland(数学数据科学教授职位,苏黎世联邦理工学院,拉沃斯,瑞士) LTS4 laboratory, EPFL, Lausanne, Switzerland(LTS4实验室,苏黎世联邦理工学院,拉沃斯,瑞士) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University, China(人工智能与高级计算学院,西安交通大学利物浦大学,中国) Cardiology Department, Lausanne University Center Hospital, Lausanne, Switzerland(心血管科,拉沃斯大学中心医院,拉沃斯,瑞士) OLV Hospital, Aalst, Belgium(OLV医院,阿尔斯特,比利时) AI Center, EPFL, Lausanne, Switzerland(人工智能中心,苏黎世联邦理工学院,拉沃斯,瑞士)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

AI总结 PINS-CAD通过基于物理的自监督学习框架,利用合成数据预训练图神经网络,提升样本效率并生成具有生理意义的预测模型,用于冠状动脉疾病的预测和预防。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03219 2025-12-04 cs.LG 57%

Perch 2.0 transfers 'whale' to underwater tasks

Perch 2.0将'鲸'转移到水下任务

Andrea Burns, Lauren Harrell, Bart van Merriënboer, Vincent Dumoulin, Jenny Hamer, Tom Denton

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 预训练与数据 :foundation model(abstract);分类 cs.LG

AI总结 Perch 2.0通过少样本迁移学习在海洋哺乳动物分类中表现出色,优于其他预训练生物声学模型。

Comments 8 pages, 3 figures, 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: AI for Non-Human Animal Communication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03158 2025-12-04 cs.LG q-bio.GN 57%

Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing

对比学习在污水基因组测序中的变异检测中的深度学习

Adele Chinda, Richmond Azumah, Hemanth Demakethepalli Venkateswara

机构 * Georgia State University(佐治亚州立大学)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.LG

AI总结 本文提出了一种基于VQ-VAE的无监督方法,用于污水基因组测序中的变异检测,通过对比学习提高变异区分能力,实现高准确率的变异识别。

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03918 2025-12-04 cs.CV 50%

UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework

UniMo:基于自回归框架统一2D视频与3D人体运动

Youxin Pang, Yong Zhang, Ruizhi Shao, Xiang Deng, Feng Gao, Xu Xiaoming, Xiaoming Wei, Yebin Liu

机构 * Tsinghua University(清华大学) Meituan(美团)

专题命中 预训练与数据 :LLM(abstract)

AI总结 UniMo通过自回归框架统一2D视频与3D人体运动,实现同时生成与理解,提升多模态联合建模能力。

Comments https://carlyx.github.io/UniMo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11056 2025-12-04 cs.CV 50%

Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization

流到模式:用于最新图像标记化的模式寻求扩散自编码器

Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 预训练与数据 :post-training(abstract)

AI总结 FlowMo是一种基于Transformer的扩散自编码器,通过模式匹配和模式寻求阶段实现图像标记化的新SOTA,无需卷积、对抗损失等。

Comments ICCV 2025, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 指令微调 9 篇

2512.03976 2025-12-04 cs.CL 91%

Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study

将大型语言模型适应于低资源藏语:一种两阶段持续和监督微调研究

Lifeng Chen, Ryan Lai, Tianming Liu

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);foundation model(abstract);pretraining(abstract)

AI总结 本研究通过两阶段方法将Qwen2.5-3B适应到藏语,通过持续预训练和监督微调提升翻译质量,实现低资源语言的模型适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02660 2025-12-04 cs.CL cs.LG 86%

How to Train Long-Context Language Models (Effectively)

如何有效训练长上下文语言模型

Tianyu Gao, Alexander Wettig, Howard Yen, Danqi Chen

机构 * Princeton Language and Intelligence(普林斯顿语言与智能)

专题命中 指令微调 :language model(title,abstract);instruction tuning(abstract);SFT(abstract);分类 cs.CL、cs.LG

AI总结 ProLong-8B 通过有效利用长上下文数据,实现了在长上下文任务上的卓越性能,尽管训练数据量仅为 Llama-3.1-8B-Instruct 的 5%。

Comments Accepted to ACL 2025. Our code, data, and models are available at https://github.com/princeton-nlp/ProLong

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03197 2025-12-04 cs.CL cs.AI 85%

InvertiTune: High-Quality Data Synthesis for Cost-Effective Single-Shot Text-to-Knowledge Graph Generation

InvertiTune:用于低成本单次文本到知识图谱生成的高质量数据合成

Faezeh Faez, Marzieh S. Tahaei, Yaochen Hu, Ali Pourranjbar, Mahdi Biparva, Mark Coates, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Autodesk Ascend Team, Huawei Technologies(华为 Ascend 团队) McGill University(麦吉尔大学)

专题命中 指令微调 :LLM(abstract);large language model(abstract);language model(abstract);SFT(abstract)

AI总结 InvertiTune通过高质量数据合成提升单次文本到知识图谱生成的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03418 2025-12-04 cs.CV cs.HC 85%

YOLOA: Real-Time Affordance Detection via LLM Adapter

YOLOA:通过LLM适配器实现实时的可及性检测

Yuqi Ji, Junjie Ke, Lihuo He, Jun Liu, Kaifan Zhang, Yu-Kun Lai, Guiguang Ding, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University(电子工程学院,西安电子科技大学) School of Software, Tsinghua University(软件学院,清华大学) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学)

专题命中 指令微调 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 YOLOA通过LLM适配器实现实时可及性检测,结合对象检测与可及性学习,取得高精度与高效率的平衡。

Comments 13 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04044 2025-12-04 cs.LG cs.AI cs.CR 81%

MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking

MarkTune: 改善开放权重语言模型水印中的质量-可检测性权衡

Yizhou Zhao, Zhiwei Steven Wu, Adam Block

机构 * University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学) Columbia University(哥伦比亚大学)

专题命中 指令微调 :LLM(title);language model(abstract);分类 cs.AI、cs.LG

AI总结 MarkTune通过理论框架提升开放权重语言模型中质量与可检测性的平衡,优于GaussMark并保持生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03499 2025-12-04 cs.CV cs.AI cs.CL 81%

NAS-LoRA: Empowering Parameter-Efficient Fine-Tuning for Visual Foundation Models with Searchable Adaptation

NAS-LoRA: 通过可搜索适应增强视觉基础模型的参数高效微调

Renqi Chen, Haoyang Su, Shixiang Tang

专题命中 指令微调 :foundation model(title,abstract);分类 cs.CL、cs.AI

AI总结 NAS-LoRA 通过引入可搜索适应的神经网络架构搜索块,提升视觉基础模型在特定领域中的参数高效微调性能,减少训练成本24.14%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03463 2025-12-04 cs.CV cs.AI cs.CL 81%

Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models

文本打印图像:为以文本为中心的大型视觉-语言模型训练弥合图像-文本模态差距

Shojiro Yamabe, Futa Waseda, Daiki Shiono, Tsubasa Takahashi

机构 * Turing Inc.(图灵公司) Institute of Science Tokyo(东京科学研究院) The University of Tokyo(东京大学) Tohoku University(东北大学)

专题命中 指令微调 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出文本打印图像(TPI)技术,通过生成合成图像弥合图像-文本模态差距,提升以文本为中心的大型视觉-语言模型训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03746 2025-12-04 cs.CV cs.CL 77%

Thinking with Programming Vision: Towards a Unified View for Thinking with Images

通过编程视觉思考:迈向图像思考的统一视角

Zirun Guo, Minjie Hong, Feng Zhang, Kai Jia, Tao Jin

机构 * Zhejiang University(浙江大学) ByteDance BandAI(字节跳动BandAI)

专题命中 指令微调 :large language model(abstract);language model(abstract);SFT(abstract);分类 cs.CL

AI总结 CodeVision提出一种灵活的代码作为工具框架,通过两阶段训练提升模型对图像变化和多工具推理的鲁棒性,实验显示显著提升性能并促进新兴能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03087 2025-12-04 cs.MM cs.AI 57%

When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI

当有害内容被伪装时:通过CamHarmTI揭示LVLMs的感知失败

Yanhui Li, Qi Zhou, Zhihong Xu, Huizhong Guo, Wenhai Wang, Dongxia Wang

机构 * Zhejiang University(浙江大学)

专题命中 指令微调 :language model(abstract);分类 cs.AI

AI总结 本文提出CamHarmTI基准,揭示LVLMs在识别伪装有害内容时的感知不足,并通过微调提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 后训练与偏好优化 7 篇

2503.18929 2025-12-04 cs.LG 90%

Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training

轨迹平衡与异步性:解耦探索与学习以实现快速、可扩展的LLM后训练

Brian Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan Obando-Ceron, Yoshua Bengio, Bhavya Kailkhura

机构 * Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Mila – Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) KAIST(韩国科学技术院) CIFAR Fellow

专题命中 后训练与偏好优化 :LLM(title,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

AI总结 TBA通过解耦探索与学习,提升LLM后训练的速度和性能,适用于多种任务并支持大规模数据生成。

Comments NeurIPS 2025; 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03882 2025-12-04 cs.LG 88%

Automatic Attack Discovery for Few-Shot Class-Incremental Learning via Large Language Models

通过大语言模型实现少样本类增量学习的自动攻击发现

Haidong Kang, Wei Wu, Hanling Wang

机构 * School of Computer and Communication Engineering(计算机与通信工程学院) School of Information and Communication(信息与通信学院) University of Electronic Science and Technology of China(电子科学与技术大学) Department of Strategic and Advanced Interdisciplinary Research(战略与先进跨学科研究部) Pengcheng Laboratory(鹏城实验室)

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 本文提出ACraft方法,利用大语言模型自动发现针对少样本类增量学习的最优攻击方法,有效提升攻击性能并降低成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23316 2025-12-04 cs.CL 85%

Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO

近端化偏好优化用于多样化反馈类型:对DPO的分解视角

Kaiyang Guo, Yinchuan Li, Zhitang Chen

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出PRO方法,通过分解DPO损失并恢复完整正则化项,解决似然不足确定性问题,提升对多样化反馈类型的适应能力。

Comments NeurIPS'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03400 2025-12-04 cs.LG cs.AI 84%

Better World Models Can Lead to Better Post-Training Performance

更好的世界模型可以导致更好的训练后性能

Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie, Andrew Lee

机构 * University of Michigan(密歇根大学) Princeton University(普林斯顿大学) Harvard University(哈佛大学)

专题命中 后训练与偏好优化 :post-training(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过比较不同世界建模策略,发现显式建模能提升Transformer的状态表示质量,从而增强强化学习后训练的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06424 2025-12-04 cs.CV 78%

Margin-aware Preference Optimization for Aligning Diffusion Models without Reference

基于边界的偏好优化:无需参考的扩散模型对齐

Jiwoo Hong, Sayak Paul, Noah Lee, Kashif Rasul, James Thorne, Jongheon Jeong

专题命中 后训练与偏好优化 :preference optimization(title,abstract)

AI总结 本文提出MaPO,一种无需参考的扩散模型对齐方法,通过优化偏好输出与非偏好输出之间的边界,提升T2I任务的适应性能,减少训练时间并优于现有方法。

Comments Accepted to AAAI 2026 Main Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏