arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Machine Learning · 会议 · Machine Learning

2026-05-13 至 2026-05-13 共收录 43
2605.12494 2026-05-13 cs.CV

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

重新审视光度歧义以实现高精度高斯-散射表面重建

Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xiaohan Yu, Lin Gu, Gim Hee Lee

机构 * School of Computer Science Engineering, State Key Laboratory of Complex Critical \& Software Environment, Jiangxi Research Institute, Beihang University State Key Laboratory of Virtual Reality Technology Macquarie University Tohoku University School of Computing, National University of Singapore

AI总结 本文提出AmbiSuR框架,通过高斯散射内在解决方案解决光度歧义问题,提升3D表面重建性能,实验显示在多种挑战场景中表现优异。

Comments Accepted at ICML 2026. Project page: https://fictionarry.github.io/AmbiSuR-Proj/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12389 2026-05-13 cs.CV cs.AI cs.LG

SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation

SEMIR: 图上语义少数诱导表示学习用于视觉分割

Luke James Miller, Yugyung Lee

机构 * Department of Computing, Analytics(计算、分析与数学系) University of Missouri-Kansas City, Kansas City, United States(密苏里大学-堪萨斯城分校)

AI总结 SEMIR通过学习拓扑保持的潜在图表示,解决大规模图像中小结构分割中的计算限制和类别不平衡问题,提升边界证据的准确性。

Comments 20 pages, 3 figures. Accepted at ICML 2026. Includes appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12258 2026-05-13 cs.LG

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

指令透镜分数:您的指令为多模态大语言模型提供了一个强大的对象幻觉检测器

Runhe Lai, Xinhua Lu, Yanqi Wu, Jinlun Ye, Weijiang Yu, Ruixuan Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(中山大学计算机科学与工程学院,广州,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) Key Laboratory of Machine Intelligence and Advanced Computing, MOE, Guangzhou, China(机器智能与高级计算关键实验室,教育部,广州,中国)

AI总结 本文提出InsLen,通过结合校准局部分数和上下文一致性分数,有效检测多模态大语言模型中的对象幻觉,无需额外训练或辅助模型。

Comments Accepted by ICML-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12167 2026-05-13 cs.RO cs.CV

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

从想象的未来到可执行的动作:用于机器人操作的潜在动作混合

Yajie Li, Bozhou Zhang, Chun Gu, Zipei Ma, Jiahui Zhang, Jiankang Deng, Xiatian Zhu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Shanghai Innovation Institute(上海创新研究院) Imperial College London(伦敦帝国理工学院) University of Surrey(萨里大学)

AI总结 本文提出MoLA,一种面向控制的接口,将想象的未来视频转化为可执行的表示,通过混合预训练的逆动力学模型,提升机器人操作的稳定性和通用性。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12046 2026-05-13 quant-ph cs.AI cs.LG

Rethink the Role of Neural Decoders in Quantum Error Correction

重新审视神经解码器在量子错误校正中的作用

Ge Yan, Shanchuan Li, Yuxuan Du

机构 * College of Computing Data Science, Nanyang Technological University, Singapore 639798, Singapore Department of Electrical Engineering Computer Science, Tokyo University of Agriculture \& Technology, Koganei, Tokyo, 184-8588, Japan School of Physical Mathematical Sciences, Nanyang Technological University, Singapore 639798, Singapore

AI总结 本文研究了神经解码器在表面码解码中的应用,探讨了在准确性与延迟约束下,通过架构重设计和压缩管道提升FPGA部署性能的方法,揭示了数据规模、归纳偏置和INT4量化对解码性能的影响。

Comments Accepted to ICML 2026; 33 Pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11931 2026-05-13 cs.CV

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

学会思考:通过视觉感知的自我改进训练提升多模态推理

Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence(计算机学院、多媒体软件国家工程研究中心、人工智能研究院) Hubei Key Laboratory of Multimedia(湖北多媒体重点实验室) Network Communication Engineering, Wuhan University, China(网络通信工程、武汉大学,中国) The University of Sydney, Australia(悉尼大学,澳大利亚) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

AI总结 本文提出VISTA框架,通过视觉感知的自我改进训练提升多模态推理能力,解决数据不平衡和语言先验偏差问题,实验显示在多种训练场景下提升性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11889 2026-05-13 cs.LG cs.AI

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

在贝叶斯学习中激励诚实与协作公平性

Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu, Patrick Jaillet, Bryan Kian Hsiang Low

机构 * Department of Computer Science, National University of Singapore, Singapore(新加坡国立大学计算机科学系) Research (A STAR), Singapore(新加坡A*STAR研究) Department of Electrical Engineering(电气工程系) Computer Science, Massachusetts Institute of Technology, USA(美国麻省理工学院计算机科学系)

AI总结 本文提出首个确保协作公平与激励诚实的机制,结合公平性保障的半值和基于验证集的诚实数据估值函数,通过理论分析和实验证实其有效性。

Comments Accepted to the 43rd International Conference on Machine Learning (ICML-26) as a Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11864 2026-05-13 cs.IR cs.AI cs.CV cs.MM

Very Efficient Listwise Multimodal Reranking for Long Documents

非常高效的长文档多模态重排序方法

Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh

机构 * Magellan Technology Research Institute (MTRI)(马杰拉技术研究院(MTRI))

AI总结 本文提出ZipRerank,通过轻量级查询-图像早期交互机制和单次前向传递消除自回归解码,实现高效多模态重排序,实验表明其在MMDocIR基准上性能优异且显著降低LLM推理延迟。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10684 2026-05-13 cs.LG cs.AI

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

数据Shapley在数据选择中不如随机?请问NASH

Xiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang, Nancy F. Chen, Bryan Kian Hsiang Low

机构 * Department of Computer Science, National University of Singapore, Singapore(新加坡国立大学计算机科学系) Research (A STAR), Singapore(新加坡科技研究局)

AI总结 本文提出NASH框架,通过分解目标函数并非线性聚合Shapley信息组件,提升数据选择效果,同时保持低运行成本。

Comments Accepted to the 43rd International Conference on Machine Learning (ICML-26) as a Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05680 2026-05-13 cs.CV

MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

MotionGRPO: 克服基于GRPO的自主运动恢复中组内多样性低的问题

Nanjie Yao, Junlong Ren, Wenhao Shen, Hao Wang

机构 * The Hong Kong University of Science(香港科学与技术大学) Nanyang Technological University, Singapore(南洋理工大学)

AI总结 本文提出MotionGRPO框架,通过强化学习后训练注入细粒度指导,解决自主运动恢复中组内多样性低的问题,提升局部关节精度和全局视觉合理性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02973 2026-05-13 cs.LG cs.AI

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges

结构化扩散桥:去噪扩散桥的归纳偏置

Eitan Kosman, Gabriele Serussi, Chaim Baskin

机构 * Ben-Gurion University of the Negev(贝纳亚克大学)

AI总结 本文提出了一种扩散桥框架,通过对可行解空间的刻画和对齐约束限制,实现模态翻译任务,展示了在无配对、半配对和配对场景下的稳定性能,尤其在降低配对要求的同时保持高质量。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14322 2026-05-13 stat.ML cs.LG

Doubly Outlier-Robust Online Infinite Hidden Markov Model

双重异常鲁棒在线无限隐马尔可夫模型

Horace Yiu, Leandro Sánchez-Betancourt, Álvaro Cartea, Gerardo Duran-Martin

机构 * Oxford-Man Institute of Quantitative Finance(牛津量化金融研究所) Mathematical Institute, University of Oxford(牛津大学数学研究所)

AI总结 本文提出BR-iHMM,通过引入两个可调参数平衡适应性与鲁棒性,在限价订单数据、小时电力需求和高维线性系统中,将一阶预测误差降低67%,并提供有界后验影响函数的理论保证。

Comments 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04476 2026-05-13 cs.CV

Vision-aligned Latent Reasoning for Multi-modal Large Language Model

多模态大语言模型中的视觉对齐潜在推理

Byungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho, Jinwoo Shin

机构 * Byungwoo Jeon Yoonwoo Jeong Hyunseok Lee Minsu Cho Jinwoo Shin

AI总结 本文提出视觉对齐潜在推理框架,通过动态生成视觉对齐的潜在标记,提升多模态大语言模型在长上下文理解和精确视觉感知任务中的表现。

Comments Published as conference proceeding for ICML 2026. Last two authors advised equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02406 2026-05-13 stat.ML cs.LG

Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

可证明的数据驱动多超参数调优与结构化损失函数

Tung Quoc Le, Anh Tuan Nguyen, Viet Anh Nguyen

机构 * Université Grenoble Alpes, LJK, CNRS, Grenoble INP(格拉诺布尔大学,LJK,CNRS,格拉诺布尔INP) Carnegie Mellon University, Machine Learning Department(卡内基梅隆大学,机器学习系) Chinese University of Hong Kong, Department of Systems Engineering and Engineering Management(香港中文大学,系统工程与工程管理系)

AI总结 本文提出首个通用框架,为数据驱动环境下多维超参数调优提供泛化保证,结合实代数几何工具,改进了半代数函数类的泛化保证,并扩展至验证损失下的超参数调优,推导出更优的界。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02282 2026-05-13 cs.LG

MoLF: Mixture-of-Latent-Flow for Pan-Cancer Spatial Gene Expression Prediction from Histology

MoLF:基于组织学的跨癌症空间基因表达预测的潜在流混合

Susu Hu, Stefanie Speidel

机构 * Translational Surgical Oncology, National Center for Tumor Diseases (NCT/UCC) Dresden, Germany Faculty of Medicine University Hospital Carl Gustav Carus, Dresden University of Technology German Cancer Research Center (DKFZ), Heidelberg, Germany

AI总结 MoLF通过条件流匹配目标,利用混合专家架构实现跨癌症组织基因表达预测,优于现有方法并在跨物种数据中表现出零样本泛化能力。

Comments Accepted at Proceedings 43rd International Conference on Machine Learning, Seoul, South Korea

Journal ref Proceedings 43rd International Conference on Machine Learning 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00767 2026-05-13 cs.LG cs.AI

BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking

BLOCK-EM:通过潜在阻塞防止对齐问题

Muhammed Ustaomeroglu, Guannan Qu

AI总结 本文提出BLOCK-EM方法,通过阻塞内部特征减少语言模型在微调中出现的对齐问题,实验显示在六个领域中阻塞固定特征可使对齐问题减少95%,且不影响模型质量与任务表现。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00297 2026-05-13 cs.LG

From Observations to States: Latent Time Series Forecasting

从观测到状态:潜在时间序列预测

Jie Yang, Yifan Hu, Yuante Li, Kexin Zhang, Kaize Ding, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) Northwestern University(西北大学)

AI总结 本文提出LatentTSF方法,通过将时间序列预测从观测回归转向潜在状态预测,解决潜在混沌问题,提升预测准确性和表示质量。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03853 2026-05-13 cs.CV

UGround: Towards Unified Visual Grounding with Unrolled Transformers

UGround: 向统一视觉接地迈进的展开变换

Rui Qian, Xin Yin, Chuanhang Deng, Zhiyuan Peng, Jian Xiong, Wei Zhai, Dejing Dou

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Zhejiang University(浙江大学)

AI总结 UGround通过引入Policy-Prompted Masking机制,动态选择展开变换层作为掩码提示,解决传统方法依赖固定最后一层和隐式投影的问题,统一了从传统参照表达分割到新提出的推理分割等多种视觉接地任务。

Comments This work has been accepted to ICML 2026, please refer to https://github.com/rui-qian/UGround

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03206 2026-05-13 cs.AI cs.CL

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

连续离散扩散:使你的扩散语言模型成为潜在推理器

Cai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang, Chubin Zhang, Muhan Zhang, Lester Mackey, Tommi Jaakkola, Stephen Bates, Dinghuai Zhang

机构 * Massachusetts Institute of Technology(麻省理工学院) Microsoft Research(微软研究院) Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥)) Peking University(北京大学) Tsinghua University(清华大学)

AI总结 本文提出CCDD模型,结合连续和离散空间,提升扩散语言模型的表达能力和训练效果,通过联合多模态扩散过程实现高质量的生成与推理。

Comments 29 pages. Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25239 2026-05-13 cs.AI cs.CL cs.LG

A Formal Comparison Between Chain of Thought and Latent Thought

链式思维与潜在思维的正式比较

Kevin Xu, Issei Sato

机构 * Department of Computer Science, The University of Tokyo, Japan(东京大学计算机科学系)

AI总结 本文对比了链式思维与潜在思维,发现潜在思维在并行计算上更高效,而链式思维能通过随机解码实现近似计数与采样。

Comments Camera-ready version for ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20754 2026-05-13 stat.ML cs.LG

Stationary MMD Points

静止MMD点

Zonghao Chen, Toni Karvonen, Heishiro Kanagawa, François-Xavier Briol, Chris. J. Oates

机构 * University College London, London, UK(伦敦大学学院) Lappeenranta--Lahti University of Technology LUT, Lappeenranta, Finland(拉佩伦塔-拉赫蒂技术大学LUT) Newcastle University, Newcastle upon Tyne, UK(纽卡斯尔大学) The Alan Turing Institute, London, UK(艾伦·图灵研究所)

AI总结 本文研究了静止MMD点的数值积分误差收敛性,证明其比MMD更快收敛,并提出基于MMD梯度流的方法计算静止点。

Journal ref International Conference on Machine Learning, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02072 2026-05-13 cs.CL cs.AI

Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs

表达怀疑——概率世界建模不应基于token logprobs

Eitan Wagner, Omri Abend

机构 * Eitan Wagner Omri Abend

AI总结 本文探讨了大型语言模型作为概率估计器在世界概率估计中的应用问题,指出基于token logprobs的局限性,并提倡第二阶预测方法以提高概率合理性。

Comments Accepted to ICML 2026 (position track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18594 2026-05-13 cs.LG stat.ML

Local and Mixing-Based Algorithms for Gaussian Graphical Model Selection from Glauber Dynamics

局部与混合基算法用于从Glauber动力学中选择高斯图模型

Vignesh Tirukkonda, Anirudh Rayas, Gautam Dasarathy

机构 * Arizona State University(亚利桑那州立大学)

AI总结 本文研究了从Glauber动力学单轨迹中学习图结构,提出局部边测试估计器和预热/稀释方法,通过总变差界证明子采样轨迹接近独立样本,提供有限样本恢复保证和信息论下限。

Comments Major revision. Corrects the earlier local ratio-estimator analysis by replacing it with a local product estimator; adds a burn-in/thinning estimator based on total-variation decoupling for Gaussian Gibbs samplers; strengthens the lower bounds; adds experiments; and compares with the related ICML 2026 work of Shen, Wu, Majid, and Moitra

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11744 2026-05-13 cs.CL cs.LG

Training-Inference Consistent Segmented Execution for Long-Context LLMs

长上下文LLM的训练-推理一致分段执行

Xianpeng Shang, Jiang Li, Zehua Duo, Qianyi Cai, Xiangdong Su

机构 * College of Computer Science, Inner Mongolia University, Hohhot 010021, China National \& Local Joint Engineering Research Center of Intelligent Information Processing Technology for Mongolian, Hohhot 010021, China Inner Mongolia Key Laboratory of Multilingual Artificial Intelligence Technology, Hohhot 010021, China Thrust of Artificial Intelligence, The Hong Kong University of Science

AI总结 本文提出训练-推理一致的分段生成框架,通过统一训练和推理的分段前向执行语义,实现长上下文生成的高效与可扩展性。

Comments Accepted by ICML 2026. 19 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11712 2026-05-13 cs.AI

Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance

迈向稳定的值对齐:引入独立模块以实现一致的价值引导

Wenhao Chen, Sirui Sun, Shengyuan Bai, Guojie Song

机构 * School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院) Yuanpei College, Peking University(北京大学元培学院) State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室)

AI总结 本文提出SVGT模型,通过独立的价值模块解决大语言模型与人类价值观对齐的问题,采用独立价值建模和显式行为引导,提升价值表达的稳定性,实验显示有效降低有害评分。

Comments Accepted to ICML 2026 (Spotlight). 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11711 2026-05-13 cs.LG cs.AI

Debiased Model-based Representations for Sample-efficient Continuous Control

去偏的基于模型的表示用于样本高效的连续控制

Jiafei Lyu, Zichuan Lin, Scott Fujimoto, Kai Yang, Yangkun Chen, Saiyong Yang, Zongqing Lu, Deheng Ye

机构 * Tencent Hunyuan(腾讯文言) McGill University(麦吉尔大学) School of Computer Science, Peking University(北京大学计算机学院)

AI总结 本文提出DR.Q算法,通过最大化当前状态-动作对表示与下一状态之间的互信息,并采用衰减优先经验回放,以减少表示和actor-critic学习中的偏差,从而在连续控制任务中取得优于现有方法的性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11622 2026-05-13 cs.CV

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

RNA-FM:基于WSI的全基因组RNA-Seq预测的流匹配生成模型

Yaxuan Song, Jianan Fan, Tianyi Wang, Qiuyue Hu, Hang Chang, Heng Huang, Weidong Cai

机构 * School of Computer Science, The University of Sydney, Australia(悉尼大学计算机科学学院) Engineering Division, Lawrence Berkeley National Lab, USA(伯克利国家实验室工程部) Berkeley Biomedical Data Science Center, Lawrence Berkeley National Lab, USA(伯克利生物医学数据科学中心) Department of Computer Science, University of Maryland College Park, USA(马里兰大学学院市计算机科学系)

AI总结 本文提出RNA-FM模型,通过流匹配生成框架实现基于WSI的全基因组RNA-Seq预测,利用连续时间条件传输问题学习速度场,实现基因表达分布的映射,提升预测的可扩展性和生物解释性。

Comments 15 pages, 13 tables, 3 figures. Accepted by the Forty-Third International Conference on Machine Learning (ICML2026). Code is available at https://github.com/YXSong000/RNA-FM

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11387 2026-05-13 cs.LG cs.RO

Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies

多模态生成策略细调中的行为模式发现

Alberta Longhini, David Emukpere, Jean-Michel Renders, Seungsu Kim

机构 * Naver Labs Europe(纳维尔实验室欧洲分部) Department of Computer Science, Stanford University(斯坦福大学计算机科学系)

AI总结 本文提出一种无监督的行为模式发现框架,用于在强化学习细调多模态生成策略时保持行为多样性,通过互信息作为内在奖励提升任务成功率。

Journal ref International Conference on Machine Learning, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11161 2026-05-13 cs.LG cs.AI

Interpretability Can Be Actionable

可解释性可以是可操作的

Hadas Orgad, Fazl Barez, Tal Haklay, Isabelle Lee, Marius Mosbach, Anja Reusch, Naomi Saphra, Byron Wallace, Sarah Wiegreffe, Eric Wong, Ian Tenney, Mor Geva

机构 * Kempner Institute at Harvard University(哈佛大学凯默纳研究所) University of Southern California(美国南加州大学) Mila – Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学) Google DeepMind(谷歌DeepMind) Tel Aviv University(特拉维夫大学) University of Pennsylvania(宾夕法尼亚大学) University of Maryland(马里兰大学) University of Oxford(牛津大学) Northeastern University(东北大学) Boston University(波士顿大学)

AI总结 本文探讨了可解释性研究的核心问题,提出应以可操作性作为评价标准,通过具体性和验证性两个维度分析阻碍实际应用的障碍,并提出五个领域和评估框架。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10815 2026-05-13 cs.AI eess.AS

Probing Cross-modal Information Hubs in Audio-Visual LLMs

探测音频-视觉大语言模型中的跨模态信息枢纽

Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim, Joon Son Chung

机构 * Department of Electrical Engineering, Korea Advanced Institute of Science The Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University, Seoul, Republic of Korea

AI总结 本文研究了音频-视觉大语言模型中音频与视觉模态间的跨模态信息流动,发现信息主要存储在sink tokens中,并提出一种无需训练的hallucination缓解方法。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏