arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

NeurIPS

Conference on Neural Information Processing Systems · 会议 · Machine Learning

共收录 17318
2502.05795 2026-02-24 cs.LG cs.AI

The Curse of Depth in Large Language Models

深度在大语言模型中的诅咒

Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu

机构 * Westlake University(西湖大学) Emory University(埃默里大学) Dalian University of Technology(大连理工大学) University of Surrey(萨里大学) University of Oxford(牛津大学)

AI总结 本文提出LayerNorm Scaling方法,通过缩放输出方差缓解深度诅咒,提升大语言模型的预训练和微调性能。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14722 2026-02-24 cs.CY econ.GN q-fin.EC

When AI Democratizes Exploitation: LLM-Assisted Strategic Manipulation of Fair Division Algorithms

当AI普及操纵:LLM辅助的公平分配算法战略操纵

Priyanka Verma, Balagopal Unnikrishnan

AI总结 本文研究了LLM如何通过简化策略知识获取,使用户能操纵公平分配算法,揭示了协调偏好操纵在资源分配中的应用及影响。

Comments accepted at NeurIPS 2025 workshop on Algorithmic Collective Action

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01485 2026-02-24 cs.LG cs.AI cs.NA math.NA stat.ML

Robust low-rank training via approximate orthonormal constraints

通过近似正交约束实现鲁棒低秩训练

Dayana Savostianova, Emanuele Zangrando, Gianluca Ceruti, Francesco Tudisco

AI总结 本文提出了一种鲁棒低秩训练算法,通过近似正交约束在保持模型精度的同时提升对抗鲁棒性。

Journal ref Proceedings NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13887 2026-02-23 eess.IV cs.AI cs.LG stat.ML

Incomplete Multi-view Clustering via Hierarchical Semantic Alignment and Cooperative Completion

不完整多视图聚类 via 层次语义对齐与合作完成

Xiaojian Ding, Lin Zhao, Xian Li, Xiaoying Zhu

机构 * School of Computer and Artificial Intelligence, Nanjing University of Finance and Economics(计算机与人工智能学院,南京财经大学)

AI总结 本文提出HSACC框架,通过层次语义对齐与动态权重分配解决不完整多视图聚类问题,实现鲁棒融合与协同学习。

Comments 13 pages, conference paper. Accepted to the Thirty-ninth Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02922 2026-02-23 cs.LG cs.AI cs.CL

Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning

克服交叉编码器中稀疏性伪影以解释聊天微调

Julian Minder, Clément Dumas, Caden Juang, Bilal Chugtai, Neel Nanda

AI总结 本文提出Latent Scaling和BatchTopK损失改进交叉编码器,以更准确地识别微调过程中出现的概念,提升模型差异分析的解释能力。

Comments 51 pages, 33 figures, Accepted at 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18322 2026-02-23 cs.LG stat.ML

Uncertainty Estimation by Flexible Evidential Deep Learning

通过灵活可信深度学习进行不确定性估计

Taeseong Yoon, Heeyoung Kim

机构 * Department of Industrial and Systems Engineering, KAIST(工业与系统工程系,韩国科学技术院)

AI总结 本文提出灵活可信深度学习(F-EDL),通过预测灵活的Dirichlet分布来提升不确定性量化在复杂和挑战性场景中的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18150 2026-02-23 cs.LG q-bio.QM stat.ML

Generative Distribution Embeddings: Lifting autoencoders to the space of distributions for multiscale representation learning

生成分布嵌入:将自编码器提升到分布空间以进行多尺度表示学习

Nic Fishman, Gokul Gowri, Peng Yin, Jonathan Gootenberg, Omar Abudayyeh

AI总结 GDE通过提升自编码器到分布空间,实现多尺度表示学习,应用于计算生物学多个关键问题,展现更强性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17497 2026-02-20 cs.LG

Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models

回顾上下文学习用于大语言模型中的时间信用分配

Wen-Tse Chen, Jiayu Chen, Fahim Tajwar, Hao Zhu, Xintong Duan, Ruslan Salakhutdinov, Jeff Schneider

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of Hong Kong(香港大学) Stanford University(斯坦福大学)

AI总结 本文提出利用大语言模型进行回顾上下文学习,以提高强化学习中时间信用分配的样本效率和泛化能力。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02819 2026-02-20 cs.CL cs.AI cs.LG

ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization

ReplaceMe: 通过深度剪枝和Transformer块线性化实现网络简化

Dmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip, Valentin Malykh, Stamatios Lefkimmiatis, Nikos Komodakis, Sergey Zagoruyko

机构 * MWS AI, ITMO University(MWS AI,ITMO大学) MWS AI MWS AI, ITMO University, IITU University(MWS AI,ITMO大学,IITU大学) University of Crete, IACM-Forth, Archimedes Athena RC(希腊克里特大学,IACM-第四研究机构,Archimedes Athena RC)

AI总结 ReplaceMe通过深度剪枝和Transformer块线性化实现高效网络简化,无需额外训练即可实现高达25%的剪枝率并保持90%性能。

Comments This work was accepted and presented at NeurIPS 2025. Code is available at https://github.com/mts-ai/replaceme Reviews at OpenReview: https://openreview.net/forum?id=zEj1FSYCRn NeurIPS 2025 Proceedings: https://openreview.net/pdf?id=zEj1FSYCRn

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17338 2026-02-20 cs.AI cs.LG stat.ML

Capturing Individual Human Preferences with Reward Features

通过奖励特征捕捉个体人类偏好

André Barreto, Vincent Dumoulin, Yiran Mao, Mark Rowland, Nicolas Perez-Nieves, Bobak Shahriari, Yann Dauphin, Doina Precup, Hugo Larochelle

机构 * Google DeepMind(谷歌DeepMind)

AI总结 通过奖励特征捕捉个体偏好,提出自适应奖励模型架构,展示其在不同用户偏好下的有效性。

Comments Published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10361 2026-02-20 cs.CL cs.LG

Enhancing Multilingual LLM Pretraining with Model-Based Data Selection

通过基于模型的数据选择增强多语言大语言模型预训练

Bettina Messmer, Vinko Sabolčec, Martin Jaggi

机构 * EPFL(苏黎世联邦理工学院)

AI总结 本文提出了一种基于模型的数据选择框架,通过提高多语言大语言模型预训练的效果,实现了在较少训练数据下达到基线分数并提升其他基准测试的表现。

Comments NeurIPS 2025 Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04485 2026-02-19 cs.LG cs.AI math.OC

Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training

Q3R:二次加权秩正则化器用于有效的低秩训练

Ipsita Ghosh, Ethan Nguyen, Christian Kümmerle

AI总结 Q3R通过二次加权秩正则化器实现有效的低秩训练,能够在保持预测性能的同时,将权重矩阵训练到指定的低秩目标。

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08822 2026-02-19 cs.RO cs.AI

FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency

FreqPolicy: 通过频率一致性实现高效的基于流的视觉-运动策略

Yifei Su, Ning Liu, Dong Chen, Zhen Zhao, Kun Wu, Meng Li, Zhiyuan Xu, Zhengping Che, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) NLPR, MAIS, Institute of Automation of Chinese Academy of Sciences(神经语言处理实验室、人工智能研究所、中国科学院自动化研究所)

AI总结 FreqPolicy通过频率一致性约束提升基于流的视觉-运动策略的效率与质量,实现高效、高质量的动作生成。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24205 2026-02-19 cs.LG stat.ML

On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks

关于混合专家模型在结构复杂任务中的表达能力

Mingze Wang, Weinan E

机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) AI for Science Institute, Beijing, China(北京人工智能科学研究院)

AI总结 本文研究了混合专家模型在结构复杂任务中的表达能力,证明了浅层和深层MoEs在低维性和稀疏性先验下能够高效近似复杂函数,并揭示了关键架构组件的作用。

Comments 28 pages, NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15076 2026-02-18 cs.LG stat.ML

Near-Optimal Sample Complexity for Online Constrained MDPs

在线约束马尔可夫决策过程的近最优样本复杂度

Chang Liu, Yunfan Li, Lin F. Yang

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

AI总结 本文提出了一种基于模型的对偶算法,解决了在线约束马尔可夫决策过程的安全性问题,证明了在允许小规模违规的情况下,算法能以近最优样本复杂度生成安全策略。

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07860 2026-02-18 cs.CV

THUNDER: Tile-level Histopathology image UNDERstanding benchmark

THUNDER:切片级病理图像理解基准

Pierre Marza, Leo Fillioux, Sofiène Boutaj, Kunal Mahatha, Christian Desrosiers, Pablo Piantanida, Jose Dolz, Stergios Christodoulidis, Maria Vakalopoulou

AI总结 THUNDER是一个用于数字病理学基础模型的切片级基准测试,通过多种数据集和下游任务评估模型的性能、特征空间及鲁棒性。

Comments Accepted at NeurIPS 2025 Datasets and Benchmarks Track (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05316 2026-02-17 cs.LG cs.AI cs.CL

Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay

通过难度目标在线数据选择和回放提升大语言模型强化微调的数据效率

Yifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen, Zhendong Wang, Mingyuan Zhou, Huan Zhang

机构 * UIUC(伊利诺伊大学香槟分校) New York University(纽约大学) University of Texas at Austin(得克萨斯大学奥斯汀分校) Microsoft(微软)

AI总结 本文提出通过难度目标在线数据选择和回放机制提升大语言模型强化微调的数据效率,实验表明可减少62%的微调时间并保持同等性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15751 2026-02-17 cs.LG cs.AI cs.CL

Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning

稀疏MeZO:在零阶LLM微调中更少的参数带来更好的性能

Yong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng, Cho-Jui Hsieh, Yang You

机构 * National University of Singapore(新加坡国立大学) Pennsylvania State University(宾夕法尼亚州立大学) University of California, Los Angeles(加州大学洛杉矶分校)

AI总结 稀疏MeZO通过仅对部分参数应用零阶优化,提升LLM微调的性能和收敛速度。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14272 2026-02-17 cs.LG

Radial-VCReg: More Informative Representation Learning Through Radial Gaussianization

径向VCReg:通过径向高斯化获得更具信息量的表示学习

Yilun Kuang, Yash Dagade, Deep Chakraborty, Erik Learned-Miller, Randall Balestriero, Tim G. J. Rudner, Yann LeCun

机构 * New York University(纽约大学) Duke University(杜克大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Brown University(布朗大学) University of Toronto(多伦多大学)

AI总结 Radial-VCReg通过引入径向高斯化损失,提升表示学习的信息量和多样性,改进自监督学习性能。

Comments Published in the Unifying Representations in Neural Models (UniReps) and Symmetry and Geometry in Neural Representations (NeurReps) Workshops at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16051 2026-02-17 astro-ph.IM cs.LG

Graph Neural Networks for Interferometer Simulations

用于干涉仪模拟的图神经网络

Sidharth Kannan, Pooyan Goodarzi, Evangelos E. Papalexakis, Jonathan W. Richardson

机构 * College of Creative Studies University of California, Santa Barbara(加州大学圣芭芭拉分校创意研究学院) Department of Physics & Astronomy University of California, Riverside(加州大学河滨分校物理与天文学系) Department of Computer Science & Engineering University of California, Riverside(加州大学河滨分校计算机科学与工程系)

AI总结 本文提出利用图神经网络模拟LIGO干涉仪,实现高效准确的光学物理仿真,并提供基准数据集用于未来研究。

Comments 11 pages, 4 figures, Accepted and Presented to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025): AI for Science Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13697 2026-02-17 cs.CY cs.AI cs.CL

Writing in Symbiosis: Mapping Human Creative Agency in the AI Era

共生写作:人工智能时代人类创造性自主性的映射

Vivan Doshi, Mengyuan Li

机构 * Department of Computer Science University of Southern California(计算机科学系美国南加州大学)

AI总结 本文探讨了人工智能时代人类创造性自主性的共生写作模式,通过分析纵向写作数据揭示了作者与AI风格的适应性变化。

Comments Advances in Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02077 2026-02-17 cs.LG

Beyond Static Cutoffs: One-Shot Dynamic Thresholding for Diffusion Language Models

超越静态截止值:用于扩散语言模型的单次动态阈值技术

Jucheng Shen, Yeonju Ro

机构 * Rice University(里奇大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出单次动态阈值技术,用于改进扩散语言模型的解码效率,在多个基准测试中提升了准确率与吞吐量的平衡。

Comments 7 pages, NeurIPS 2025 Efficient Reasoning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04398 2026-02-17 cs.CL cs.AI cs.CR cs.LG

SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations

SECA:用于诱发LLM幻觉的语义等价且连贯的攻击

Buyun Liang, Liangzu Peng, Jinqi Luo, Darshan Thaker, Kwan Ho Ryan Chan, René Vidal

AI总结 SECA通过现实修改提示诱发LLM幻觉,提高攻击成功率并减少语义错误。

Comments Accepted at NeurIPS 2025. Code is available at https://github.com/Buyun-Liang/SECA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23519 2026-02-17 cs.CR cs.AI

ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search

ReliabilityRAG: 有效且可证明的防御方法用于基于检索的Web搜索

Zeyu Shen, Basileal Imana, Tong Wu, Chong Xiang, Prateek Mittal, Aleksandra Korolova

机构 * Department of Computer Science(计算机科学系) Princeton University(普林斯顿大学) Center for Information Technology Policy(信息政策中心) Department of Electrical and Computer Engineering(电气与计算机工程系) NVIDIA(英伟达)

AI总结 ReliabilityRAG通过利用文档可靠性信息,提供更有效且可证明鲁棒的防御方法,以增强基于检索的Web搜索系统对对抗攻击的抵御能力。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19189 2026-02-17 cs.LG stat.ML

Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules

核回归中的功能缩放定律:损失动态与学习率调度

Binghui Li, Fengling Chen, Zixun Huang, Lean Wang, Lei Wu

机构 * Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) School of Mathematical Sciences, Peking University(北京大学数学科学学院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学多媒体信息处理国家重点实验室) AI for Science Institute, Beijing(北京人工智能科学研究院)

AI总结 本文提出功能缩放定律,通过分析核回归模型中的损失动态,揭示学习率调度对训练效率的影响,并验证了WSD调度在大规模预训练中的有效性。

Comments 60 pages, accepted by NeurIPS 2025 as a spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01055 2026-02-17 cs.LG cs.AI q-bio.BM q-bio.QM

FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models

FGBench: 一个用于大语言模型中功能基团级分子属性推理的数据集和基准

Xuan Liu, Siru Ouyang, Xianrui Zhong, Jiawei Han, Huimin Zhao

机构 * Department of Chemical and Biomolecular Engineering, University of Illinois Urbana-Champaign(化学与生物分子工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Computer Science, University of Illinois Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校)

AI总结 FGBench通过构建包含功能基团级信息的数据集,提升大语言模型在分子属性推理任务中的能力。

Comments NeurIPS 2025 (Datasets and Benchmarks Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07272 2026-02-17 cs.LG

A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing

基于Cramér-von Mises统计的促进真实数据共享方法

Alex Clinton, Thomas Zeng, Yiding Chen, Xiaojin Zhu, Kirthevasan Kandasamy

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Cornell University(康奈尔大学)

AI总结 本文提出基于Cramér-von Mises统计的激励机制,促进真实数据共享,通过理论分析和实验验证其有效性。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06964 2026-02-17 cs.CL cs.LG

Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization

通过奖励加权微调实现离线强化学习用于对话优化

Subhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi, Seunghyun Yoon, Trung Bui, Anup Rao, Jayakumar Subramanian, Branislav Kveton

机构 * Adobe Research(Adobe研究院)

AI总结 本文提出通过奖励加权微调实现离线强化学习,用于对话优化,相比传统方法在奖励和语言质量上取得显著提升。

Comments Advances in Neural Information Processing Systems 38

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19712 2026-02-17 cs.LG math.PR stat.ML

On the Relation between Rectified Flows and Optimal Transport

关于校正流与最优传输关系的研究

Johannes Hertrich, Antonin Chambolle, Julie Delon

机构 * Université Paris Dauphine - PSL & Inria(巴黎-萨克雷大学) Inria(法国国家信息与自动化技术研究院) ENS Paris(巴黎高等师范学校)

AI总结 本文探讨了校正流与最优传输之间的关系,指出校正流的梯度约束并不能保证最优传输的解,并提供了反例推翻了之前的等价性结论。

Comments Accepted for NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19645 2026-02-17 cs.LG cs.AI

MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE

MoESD: 解开投机解码在加速稀疏MoE中的潜力

Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu, Weikai Mao, Xianzhi Yu, Yongpan Liu, Tianyu Zhang

机构 * Tsinghua University(清华大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 MoESD研究揭示了投机解码在加速稀疏MoE模型中的潜力,提出目标效率指标以更全面理解SD加速效果。

Comments Accepted as spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏