arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Transactions on Machine Learning Research · 期刊 · Machine Learning

至 收录 112
2401.11512 2026-07-07 cs.LG cs.AI cs.IT math.IT 版本更新

TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning

TERC:一种用于强化学习状态变量选择的转移熵冗余准则

Charles Westphal, Stephen Hailes, Mirco Musolesi

机构 * UCL Centre for Artificial Intelligence(伦敦大学学院人工智能中心)

AI总结 本文提出TERC准则,用于选择强化学习中的最优状态变量,通过信息理论方法排除冗余变量,提升推理效率,适用于多种算法和环境。

Comments 47 pages, 12 figures, accepted in TMLR (https://openreview.net/forum?id=J0ad21E0vX)

URL PDF HTML 收藏
2510.20091 2026-07-03 cs.CL cs.AI 版本更新

CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

CreativityPrism:大语言模型创造力的跨域评估框架

Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li

机构 * University of Pittsburgh(匹兹堡大学) Johns Hopkins University(约翰霍普金斯大学) University of Notre Dame(诺特丹大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能研究院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 提出CreativityPrism框架,整合发散思维、创意写作和逻辑推理三个领域的八项任务,从质量、新颖性和多样性三个维度评估LLM创造力,发现前沿模型在创意写作和逻辑推理上领先,但在发散思维上无显著优势,且各维度间相关性弱。

Comments Published in Transactions on Machine Learning Research (06/2026)

URL PDF HTML 收藏
2506.09105 2026-07-03 cs.LG cs.AI quant-ph 版本更新

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

MetaTT: 一种用于参数高效微调的全局张量列适配器

Javier Lopez-Piqueres, Pranav Deshpande, Archan Ray, Mattia J. Villani, Marco Pistoia, Niraj Kumar

机构 * Global Technology Applied Research(全球技术应用研究)

AI总结 提出MetaTT,一种基于张量列(TT)分解的适配器框架,通过共享单个TT因子化Transformer子模块,实现参数高效的多任务微调,并在单任务和多任务基准上达到竞争性性能。

Comments Accepted version to TMLR

URL PDF HTML 收藏
2408.01139 2026-07-03 cs.AI cs.CV 版本更新

Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition

使用公理谱重要性分解解释图像模型的全局扰动鲁棒性

Róisín Luo, James McDermott, Colm O'Riordan

机构 * SFI Centre for Research Training in Artificial Intelligence(SFI人工智能研究培训中心) School of Computer Science, University of Galway(Galway大学计算机科学学院)

AI总结 提出一种模型无关的全局可解释性方法I-ASIDE,基于Shapley值公理量化鲁棒与非鲁棒特征的预测能力,揭示图像模型对数据损坏和对抗攻击等扰动的鲁棒性机制。

Comments Accepted by Transactions on Machine Learning Research (TMLR 2024)

Journal ref Transactions on Machine Learning Research (TMLR), 2024; Presented at The Thirteenth International Conference on Learning Representations (ICLR 2025), Singapore

URL PDF HTML 收藏
2507.10540 2026-07-02 cs.LG 版本更新

FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data

FusionFactory: 融合多LLM日志数据中的大语言模型能力

Tao Feng, Haozhen Zhang, Zijie Lei, Pengrui Han, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Nanyang Technological University(南洋理工大学) Meta Monetization AI(Meta 变现人工智能) NVIDIA(英伟达)

AI总结 提出FusionFactory框架,通过查询级、思维级和模型级融合策略,利用多LLM日志数据提升模型性能,在14个基准测试中均优于最佳单模型。

Journal ref TMLR 2026

URL PDF HTML 收藏
2601.16398 2026-07-01 cs.CY cs.CL cs.LG 版本更新

White-Box Sensitivity Auditing with Steering Vectors

白盒敏感性审计与引导向量

Hannah Cyberey, Yangfeng Ji, David Evans

机构 * University of Virginia(弗吉尼亚大学)

AI总结 本文提出白盒敏感性审计框架,通过激活引导进行更严格的模型内部评估,用于检测大语言模型中的偏见,揭示模型对保护属性的依赖。

Comments Accepted to Transactions on Machine Learning Research (TMLR)

URL PDF HTML 收藏
2603.23867 2026-07-01 cs.LG cs.AI cs.CV 版本更新

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation

VLM能稳健推理吗?一项神经符号研究

Weixin Chen, Antonio Vergari, Han Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Edinburgh(爱丁堡大学)

AI总结 研究视觉语言模型在分布偏移下的推理稳健性,提出结合VLM概念识别与电路符号推理的神经符号方法VLC,在三个视觉演绎推理任务上实现更高的分布外准确率。

Comments TMLR 2026

URL PDF HTML 收藏
2510.17139 2026-07-01 cs.CL cs.IR 版本更新

Rethinking On-policy Optimization for Query Augmentation

重新思考查询增强的在线策略优化

Zhichao Xu, Shengyao Zhuang, Xueguang Ma, Bingsen Chen, Yijun Tian, Fengran Mo, Tao Li, Jie Cao, Vivek Srikumar

机构 * University of Utah(犹他大学) The University of Queensland(昆士兰大学) University of Waterloo(滑铁卢大学) New York University(纽约大学) University of Notre Dame(圣母大学) Université de Montréal(蒙特利尔大学) Google DeepMind(谷歌DeepMind) University of Oklahoma(俄克拉荷马大学)

AI总结 本文系统比较了基于提示和强化学习的查询增强方法,发现计算量感知下简单方法常优于RL方法,并提出混合方法OPQE,通过RL生成伪文档以最大化检索性能。

Comments TMLR camera ready version

URL PDF HTML 收藏
2511.11046 2026-07-01 cs.LG cs.AI 版本更新

Enhancing Graph Representations with Neighborhood-Contextualized Message-Passing

增强图表示:邻域上下文化的消息传递

Brian Godwin Lim, Galvin Brice Lim, Renzo Roel Tan, Irwin King, Kazushi Ikeda

机构 * Nara Institute of Science and Technology(奈良先端科学技术大学院大学) Kyoto University(京都大学) Ateneo de Manila University(马尼拉雅典耀大学) UNI-President Information Philippines Corporation(统一信息菲律宾公司) The Chinese University of Hong Kong(香港中文大学)

AI总结 提出邻域上下文化消息传递(NCMP)框架,通过整合多集邻域上下文增强GNN表达能力,并实例化为SINC-GCN,在保持高效的同时显著提升性能。

Comments Published in Transactions on Machine Learning Research

Journal ref Transactions on Machine Learning Research. (2026)

URL PDF HTML 收藏
2502.00168 2026-06-29 stat.ML cs.LG math.DG math.ST stat.TH 版本更新

Supervised Quadratic Feature Analysis: Information Geometry Approach for Dimensionality Reduction

监督二次特征分析:用于降维的信息几何方法

Daniel Herrera-Esposito, Johannes Burge

AI总结 提出监督二次特征分析(SQFA),利用Fisher-Rao距离最大化类间差异,在高斯假设下学习线性特征,实验表明SQFA在分类精度上具有竞争力。

Comments 32 pages, 11 figures

Journal ref Transactions on Machine Learning Research (2026)

URL PDF HTML 收藏
2411.07175 2026-06-29 cs.CL 版本更新

Continual Memorization of Factoids in Language Models

语言模型中事实的持续记忆

Howard Chen, Jiayi Geng, Adithya Bhaskar, Dan Friedman, Danqi Chen

机构 * Princeton Language and Intelligence (PLI), Princeton University(普林斯顿大学普林斯顿语言与智能研究所)

AI总结 提出持续记忆任务,发现语言模型在后续微调中严重遗忘事实,并提出REMIX方法(混合随机词序列或预训练语料)缓解遗忘,优于回放等方法。

Journal ref Transactions on Machine Learning Research, 2026

URL PDF HTML 收藏
2509.20008 2026-06-26 cs.LG cs.CR 版本更新

Learning Robust Penetration Testing Policies under Partial Observability: A systematic evaluation

学习部分可观测下的鲁棒渗透测试策略:系统评估

Raphael Simon, Pieter Libin, Wim Mees

机构 * Cyber Defence Lab, CISS Department Royal Military Academy(国防网络安全实验室,信息与系统科学系皇家军事学院) AI Lab, Department of Computer Science Vrije Universiteit Brussel(人工智能实验室,计算机科学系自由大学布鲁塞尔)

AI总结 针对部分可观测的渗透测试问题,系统评估了多种PPO变体(如帧堆叠、历史观测增强、LSTM/TrXL架构)在主机网络中的性能,发现历史聚合策略收敛速度提升四倍,并揭示了策略的定性差异。

Comments Published in Transactions on Machine Learning Research (TMLR) https://openreview.net/forum?id=YkUV7wfk19. 25 pages, 8 figures. Code and StochNASim environment are available at https://github.com/raphsimon/StochNASim

Journal ref Transactions on Machine Learning Research, 2026

URL PDF HTML 收藏
2604.23178 2026-06-25 cs.AI 版本更新

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

评判评判者:LLM-as-a-Judge pipelines中偏见缓解策略的系统评估

Sadman Kabir Soumik

机构 * Independent Researcher(独立研究员)

AI总结 本文系统评估了LLM-as-a-Judge pipelines中九种偏见缓解策略,发现风格偏见是最主要的偏见类型,且所有模型在扩展对上偏好简洁性,但截断控制能区分质量和长度,表明质量敏感的评估而非单纯长度偏见。

Comments 22 pages, 4 figures. Published in Transactions on Machine Learning Research (2026)

Journal ref Transactions on Machine Learning Research (2026)

URL PDF HTML 收藏
2405.17423 2026-06-25 cs.CV cs.CL 版本更新

Privacy-Aware Visual Language Models

隐私感知的视觉语言模型

Laurens Samson, Nimrod Barazani, Sennay Ghebreab, Yuki M. Asano

机构 * Socially-Intelligent Artificial Systems Group, University of Amsterdam(智能社会人工智能系统组,阿姆斯特丹大学) University of Amsterdam(阿姆斯特丹大学) Fundamental AI Lab, University of Technology Nuremberg(基础人工智能实验室,纽伦堡技术大学)

AI总结 针对视觉语言模型隐私理解不足的问题,构建高质量基准数据集PrivBench和指令微调数据集PrivTune,通过少量样本微调显著提升隐私敏感性,性能超越GPT-4。

Comments Accepted at Transactions on Machine Learning Research (TMLR)

URL PDF HTML 收藏
2502.01015 2026-06-24 cs.LG 版本更新

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic

任务向量基:一种统一且可扩展的压缩任务算术框架

Siqi Zeng, Yifei He, Meitong Liu, Weiqiu You, Yifan Hao, Yao-Hung Hubert Tsai, Makoto Yamada, Han Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Pennsylvania(宾夕法尼亚大学) Okinawa Institute of Science and Technology(冲绳科学技术大学院大学)

AI总结 提出任务向量基框架,将T个任务向量压缩为M<T个基向量,支持加法、否定等算术操作,理论保证泛化与遗忘,实验表明压缩后性能优于或持平完整集合。

Comments Published as a journal paper at TMLR in Jun 2026. 10 pages, 12 figures

URL PDF HTML 收藏
2501.18502 2026-06-24 cs.IT math.IT math.ST stat.TH 版本更新

One-Bit Distributed Mean Estimation with Unknown Variance

方差未知的单比特分布式均值估计

Ritesh Kumar, Shashank Vatedka

AI总结 针对方差未知的分布式均值估计问题,提出简单非自适应和自适应协议,证明渐近正态性并推导均方误差界,通过下界证明自适应协议的最优性。

Comments Published in Transactions on Machine Learning Research https://openreview.net/forum?id=g95C4zIEPg

Journal ref Transactions on Machine Learning Research, February 2026

URL PDF HTML 收藏
2508.05469 2026-06-23 cs.LG cs.IT math.IT 版本更新

Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes

逐步测量信息:超越情绪的AI评估

Zachary Robertson, Sanmi Koyejo

机构 * Department of Computer Science(计算机科学系)

AI总结 本文提出基于AI的评估方法,通过战略游戏与信息损失的联系,分析对抗操纵下鲁棒的机制,证明总变差距离在对抗攻击下保持多项式保证,提升信息关系提示的鲁棒性。

Comments Accepted to TMLR (2026). Updated appendix to restore the proof of Theorem 3.3

URL PDF HTML 收藏
2511.18471 2026-06-23 cs.CV 版本更新

Jacobian-Aware Posterior Sampling for Inverse Problems

Jacobian-Aware 后验采样用于逆问题

Liav Hen, Tom Tirer, Raja Giryes, Shady Abu-Hussein

机构 * Tel Aviv University, Israel(特拉维夫大学,以色列) Bar-Ilan University, Israel(巴伊兰大学,以色列) University of Cambridge, UK(剑桥大学,英国)

AI总结 提出 Jacobian-Aware 后验采样器 (JAPS),通过结合扩散去噪器的 Jacobian 先验与近端解,在无额外计算成本下提升逆问题重建质量。

Comments Accepted by TMLR 2026; Code at https://github.com/liavhen/JAPS

URL PDF HTML 收藏
2601.09166 2026-06-23 cs.LG cs.CR cs.DC 版本更新

DP-FedSOFIM: Differentially Private Federated Stochastic Optimization using Regularized Fisher Information Matrix

DP-FedSOFIM: 使用正则化Fisher信息矩阵的差分隐私联邦随机优化

Sidhant Nair, Tanmay Sen, Mrinmay Sen, Sayantan Banerjee

机构 * Department of Mechanical Engineering, Indian Institute of Technology Delhi(印度理工学院德里机械工程系) SQC & OR Unit, Indian Statistical Institute Kolkata(印度统计研究院科钦SQC与OR单位) Department of Artificial Intelligence, Indian Institute of Technology Hyderabad(印度理工学院海得拉巴人工智能系) Operations Management & Quantitative Techniques Area, Indian Institute of Management, Indore(印度管理学院印地尔运营管理和定量技术领域)

AI总结 针对差分隐私联邦学习在严格隐私预算下收敛慢的问题,提出DP-FedSOFIM方法,利用正则化Fisher信息矩阵的代理进行二阶优化,无需完整Hessian计算,在CIFAR-10和PathMNIST上实现更快收敛和更高精度。

Comments 59 pages, 8 figures, 16 tables. Accepted to TMLR

URL PDF HTML 收藏
2602.17431 2026-06-23 cs.CL cs.AI cs.LG 版本更新

Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study

长文本语言模型输出的细粒度不确定性量化:一项比较研究

Dylan Bouchard, Mohit Singh Chauhan, Viren Bajaj, David Skarbrevik

机构 * CVS Health(CVS健康公司) Wellesley, MA(马萨诸塞州韦尔士利)

AI总结 针对长文本生成,提出细粒度不确定性量化分类法,通过响应分解、单元级评分和响应级聚合三阶段区分方法,引入FactScore-STEM-Geo数据集,实验发现声明-响应蕴含评分优于复杂方法,声明级评分优于句子级,不确定性感知解码有效提升事实性。

Comments Accepted by TMLR; UQLM repository: https://github.com/cvs-health/uqlm

Journal ref Transactions on Machine Learning Research, 2026

URL PDF HTML 收藏
2504.05520 2026-06-23 cs.LG cs.CL 版本更新

Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

通过自适应课程学习的高效强化微调

Taiwei Shi, Yiyang Wu, Linxin Song, Tianyi Zhou, Jieyu Zhao

机构 * University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 提出AdaRFT方法,通过自适应课程学习动态调整训练问题难度,提升强化微调效率,在数学推理任务上训练时间减半。

Comments Published in Transactions on Machine Learning Research (TMLR). 30 pages, 8 figures, 7 tables

URL PDF HTML 收藏
2509.23729 2026-06-23 cs.CV cs.AI cs.LG eess.IV 版本更新

LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models

LUQ: 多模态大语言模型的逐层超低位量化

Shubhang Bhatnagar, Andy Xu, Kar-Han Tan, Narendra Ahuja

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, Los Angeles(加州大学洛杉矶分校) HP Inc.(惠普公司)

AI总结 提出LUQ方法,通过逐层输出激活熵衡量功能复杂度,对简单层应用超低位量化,结合多模态校准,在LLaVA-1.5和Qwen-2.5-VL上实现低于4-bit量化,内存减少40%和31%,MME性能下降<10%。

Comments Published in Transactions on Machine Learning Research (2026)

URL PDF HTML 收藏
2510.02561 2026-06-23 cs.CV cs.AI 版本更新

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback

Oracle-RLAIF:一种利用排名反馈强化学习改进多模态视频模型微调框架的方法

Derek Shi, Ruben Glatt, Christine Klymko, Shubham Mohole, Hongjun Choi, Shashank Kushwaha, Sam Sakla, Felipe Leno da Silva

机构 * Stanford University(斯坦福大学) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Microsoft(微软公司)

AI总结 提出Oracle-RLAIF框架,用通用排序器替代奖励模型,结合基于GRPO的排名损失函数GRPO_rank,实现更高效的多模态视频模型微调,在多个基准上优于现有方法。

Comments Proceedings of the 39th Annual Conference on Neural Information Processing Systems, ARLET Workshop (Aligning Reinforcement Learning Experimentalists and Theorists)

Journal ref Transactions on Machine Learning Research, Vol. 2026, June 2026

URL PDF HTML 收藏
2505.01652 2026-06-23 cs.LG cs.AI 版本更新

Causally Fair Node Classification on Non-IID Graph Data

非独立同分布图数据上的因果公平节点分类

Yucong Dai, Lu Zhang, Yaowei Hu, Susan Gauch, Yongkai Wu

机构 * Clemson University(克莱姆森大学) University of Arkansas(亚拉巴马大学) Walmart Inc.(沃尔玛公司)

AI总结 针对图数据中节点因果机制异质性违反经典因果模型假设的问题,提出基于网络结构因果模型的消息传递变分自编码器,通过计算干预分布实现因果公平节点分类,理论保证可分解性和图独立性条件,实验验证其有效缓解偏差。

Comments Accepted by TMLR: https://openreview.net/forum?id=AwptwzGld5

URL PDF HTML 收藏
2503.22934 2026-06-23 cs.LG cs.AI 版本更新

FairSAM: Fair Classification on Corrupted Image Data Through Sharpness-Aware Minimization

FairSAM: 通过锐度感知最小化实现损坏图像数据上的公平分类

Yucong Dai, Jie Ji, Xiaolong Ma, Yongkai Wu

机构 * Clemson University(克莱姆森大学)

AI总结 针对图像分类模型在数据损坏时性能下降且不公平的问题,提出FairSAM框架,将公平性策略融入锐度感知最小化,平衡鲁棒性与公平性。

Comments Accepted by TMLR: https://openreview.net/forum?id=W2QKvn57yw

URL PDF HTML 收藏
2412.00143 2026-06-23 cs.LG cs.CV 版本更新

Is Oracle Pruning the True Oracle?

Oracle剪枝真的是真正的Oracle吗?

Sicheng Feng, Keda Tao, Huan Wang

机构 * Westlake University(西湖大学) Nankai University(南开大学) ENCODE Lab, Westlake University(西湖大学ENCODE实验室)

AI总结 本文通过大规模实验(37K模型)发现,对于中等规模以上的深度学习模型,Oracle剪枝选择的权重在重训练后性能与重训练前几乎无关,质疑了Oracle剪枝作为剪枝方法基础的有效性。

Comments TMLR, Webpage: https://fscdc.github.io/Oracle-Pruning-Sanity-Check/

URL PDF HTML 收藏
2502.10239 2026-06-18 cs.LG cs.AI 版本更新

Efficient Zeroth-Order Federated Finetuning of Language Models on Resource-Constrained Devices

资源受限设备上语言模型的高效零阶联邦微调

Mohamed Aboelenien Ahmed, Kilian Pfeiffer, Ramin Khalili, Heba Khdr, Jörg Henkel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Huawei(华为) Heisenberg Research Center (Munich), Germany(海森堡研究中心(慕尼黑),德国)

AI总结 提出一种基于零阶优化的联邦微调方法,通过分块模型并分配更多扰动到后一块,复用中间激活减少前向评估次数,在保持内存和通信优势的同时将计算量降低至其他零阶方法的1/3。

Comments Published at TMLR

URL PDF HTML 收藏
2507.07574 2026-06-18 cs.CV 版本更新

Beyond the Linear Separability Ceiling: Aligning Representations in VLMs

超越线性可分上限:对齐视觉-语言模型中的表征

Enrico Vompa, Tanel Tammet, Mohit Vaishnav

机构 * Applied Artificial Intelligence Group(应用人工智能小组) Tallinn University of Technology(塔林技术大学)

AI总结 提出线性可分上限(LSC)诊断框架,发现VLM存在对齐差距,并通过对比目标重塑视觉流形,使模型在抽象组合推理任务上显著超越LSC。

Comments Accepted TMLR

URL PDF HTML 收藏
2508.20330 2026-06-18 cs.LG 版本更新

FORGE: Foundational Optimization Representations from Graph Embeddings

FORGE:基于图嵌入的基础优化表示

Zohair Shafi, Serdar Kadioglu

机构 * Khoury College of Computer Science Northeastern University(诺埃弗大学计算机科学学院) AI Center of Excellence, Fidelity Investments(富达投资人工智能卓越中心) Department of Computer Science, Brown University(布朗大学计算机科学系)

AI总结 提出FORGE框架,通过无监督预训练向量量化图自编码器学习混合整数规划实例的通用表示,无需求解器或最优解,在下游任务中提升求解器性能并超越现有方法。

Comments Published in TMLR

URL PDF HTML 收藏
2411.08821 2026-06-17 stat.ML cs.LG stat.CO 版本更新

Conditional Local Importance by Quantile Expectations

基于分位数期望的条件局部重要性

Kelvyn K. Bladen, Adele Cutler, D. Richard Cutler, Kevin R. Moon

机构 * Department of Mathematics & Statistics(数学与统计学系)

AI总结 提出模型无关的局部变量重要性方法CLIQUE,通过分位数期望捕获局部依赖关系,提升稳定性并直接适用于多类分类问题。

Comments 29 pages, 28 figures

Journal ref Transactions on Machine Learning Research (2026)

URL PDF HTML 收藏