arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-17 至 2026-02-17 共收录 101 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 21 篇

2602.14017 2026-02-17 cs.LG 78%

S2SServiceBench: A Multimodal Benchmark for Last-Mile S2S Climate Services

S2SServiceBench:一个多模态基准用于最后一公里S2S气候服务

Chenyue Li, Wen Deng, Zhuotao Sun, Mengxi Jin, Hanzhe Cui, Han Li, Shentong Li, Man Kit Yu, Ming Long Lai, Yuhao Yang, Mengqian Lu, Binhang Yuan

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Nanjing University of Information Science and Technology(南京信息工程大学) Beijing Normal University(北京师范大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 S2SServiceBench是一个多模态基准,用于评估S2S气候服务中最后一公里的可靠性,通过10种服务产品和1000多个评估项目,揭示了多模态大语言模型在不确定性下的决策推理挑战。

Comments 18 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13269 2026-02-17 cs.NI 78%

Modality-Tailored Age of Information for Multimodal Data in Edge Computing Systems

多模态数据在边缘计算系统中的模态定制信息年龄

Ying Liu, Yifan Zhang, Xinyu Wang, Chao Yang, Kandaraj Piamrat, Stephan Sigg, Zheng Changr, Yusheng Ji

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出模态定制信息年龄(MAoI)度量标准,用于多模态数据在边缘计算中的资源管理和策略优化,并设计了联合采样卸载优化算法以最小化平均 MAoI。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15011 2026-02-17 cs.HC 71%

TouchFusion: Multimodal Wristband Sensing for Ubiquitous Touch Interactions

TouchFusion: 多模态腕带传感用于无处不在的触控交互

Eric Whitmire, Evan Strasnick, Roger Boldu, Raj Sodhi, Nathan Godwin, Shiu Ng, Andre Levi, Amy Karlson, Ran Tan, Josef Faller, Emrah Adamey, Hanchuan Li, Wolf Kienzle, Hrvoje Benko

专题命中 多模态评测 :multimodal(title)

AI总结 TouchFusion通过多模态腕带传感实现无处不在的触控交互,结合多种传感器技术,支持环境和身体表面的触控检测与上下文自适应界面控制。

Comments 23 pages, 22 figures, accompanying video available at https://youtu.be/0fdCwHu7uaA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20430 2026-02-17 cs.CL cs.AI cs.CV cs.MA 67%

An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

一种具有可追溯推理的罕见病诊断代理系统

Weike Zhao, Chaoyi Wu, Yanjie Fan, Xiaoman Zhang, Pengcheng Qiu, Yuze Sun, Xiao Zhou, Yanfeng Wang, Xin Sun, Ya Zhang, Yongguo Yu, Kun Sun, Weidi Xie

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DeepRare是一种基于大型语言模型的多代理系统,通过整合40多种专用工具和最新知识源,为罕见病诊断提供决策支持,实现了透明可追溯的推理链。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05430 2026-02-17 cs.CL cs.AI cs.IR cs.LG 62%

ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering

ArtistMus: 一个全球多样、以艺术家为中心的基准,用于检索增强的音乐问答

Daeyong Kwon, SeungHeon Doh, Juhan Nam

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 ArtistMus提出一个全球多样、以艺术家为中心的基准,通过检索增强生成技术提升音乐问答的准确性和上下文推理能力。

Comments Accepted to LREC 2026. This work is an evolution of our earlier preprint arXiv:2507.23334

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01112 2026-02-17 cs.CL 57%

EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation

EmoLoom-2B:基于词典弱监督和KV-Off评估的快速基础模型筛选用于情感分类和VAD预测

Zilin Li, Weiwei Xu, Xuanbo Lu, Zheda Liu

机构 * Zheda Liu(2 刘智达)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 EmoLoom-2B通过词典弱监督和KV-Off评估,快速筛选出适用于情感分类和VAD预测的基础模型。

Comments This paper presents an initial and self-contained study of a lightweight screening pipeline for emotion-aware language modeling, intended as a reproducible baseline and system-level design reference. This latest version corrects and updates certain personal information

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13806 2026-02-17 cs.CV cs.RO 57%

Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos

具有多尺度动态的高斯序列用于从单目随意视频中进行4D重建

Can Li, Jie Gu, Jingmin Chen, Fangzhou Qiu, Lei Sun

机构 * Nankai University(南开大学) Rightly Robotics

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于多尺度动态的高斯序列方法,用于从单目随意视频中实现准确且一致的4D重建,通过多级运动组合和多模态先验约束提升重建保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13771 2026-02-17 cs.MM 57%

SRA: Semantic Relation-Aware Flowchart Question Answering

SRA: 语义关系感知的流程图问答

Xinyu Li, Bowei Zou, Yuchong Chen, Yifan Fan, Yu Hong

专题命中 多模态评测 :multi-modal(abstract);分类 cs.MM

AI总结 SRA通过利用大型语言模型检测节点间的语义关系,提升流程图问答的推理深度和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00220 2026-02-17 eess.IV cs.CV 57%

Deep learning Based Correction Algorithms for 3D Medical Reconstruction in Computed Tomography and Macroscopic Imaging

基于深度学习的3D医学重建在计算机断层扫描和宏观成像中的校正算法

Tomasz Les, Tomasz Markiewicz, Malgorzata Lorent, Miroslaw Dziekiewicz, Krzysztof Siwek

机构 * University of Technology(技术大学) Military Institute of Medicine(军事医学研究院) Institute of Tuberculosis and Lung Diseases(肺结核和肺病研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种混合两阶段配准框架,结合几何先验和深度学习,提升3D医学重建的精度和解剖真实感。

Comments 23 pages, 9 figures, submitted to Applied Sciences (MDPI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01055 2026-02-17 cs.LG cs.AI q-bio.BM q-bio.QM 57%

FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models

FGBench: 一个用于大语言模型中功能基团级分子属性推理的数据集和基准

Xuan Liu, Siru Ouyang, Xianrui Zhong, Jiawei Han, Huimin Zhao

机构 * Department of Chemical and Biomolecular Engineering, University of Illinois Urbana-Champaign(化学与生物分子工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Computer Science, University of Illinois Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 FGBench通过构建包含功能基团级信息的数据集,提升大语言模型在分子属性推理任务中的能力。

Comments NeurIPS 2025 (Datasets and Benchmarks Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19666 2026-02-17 cs.CL 57%

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

RoD-TAL:罗马尼亚驾照考试问答的基准测试

Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel

机构 * National University of Science and Technology POLITEHNICA Bucharest, Faculty of Automatic Control and Computers(波兰技术大学布加勒斯特分校) Technical University of Munich(慕尼黑技术大学) National Institute for Research & Development in Informatics - ICI Bucharest(信息研究所-布加勒斯特) Paris 1 Panthéon-Sorbonne University(巴黎1大学) University of Bucharest(布加勒斯特大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 RoD-TAL是一个用于评估大型语言模型和视觉语言模型在罗马尼亚驾照法律问答中性能的多模态基准数据集,通过文本和图像问答任务验证了领域微调和推理优化对考试通过率的影响。

Comments 41 pages, 30 figures, Accepted by the Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14061 2026-02-17 stat.CO cs.NA math.NA 50%

MPL-HMC: A Tunable Parameterized Leapfrog Framework for Robust Hamiltonian Monte Carlo

MPL-HMC:一种可调参数化Leapfrog框架用于鲁棒哈密顿蒙特卡洛

Sourabh Bhattacharya

专题命中 多模态评测 :multimodal(abstract)

AI总结 MPL-HMC通过可调参数改进HMC,实现鲁棒采样和性能提升,适用于多模分布和复杂模型。

Comments Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02410 2026-02-17 cs.LG 50%

OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data

OpenTSLM:用于多变量医学文本和时间序列数据推理的时间序列语言模型

Patrick Langer, Thomas Kaar, Max Rosenblattl, Maxwell A. Xu, Winnie Chow, Martin Maritsch, Robert Jakob, Ning Wang, Juncheng Liu, Aradhana Verma, Brian Han, Daniel Seung Kim, Henry Chubb, Scott Ceresnak, Aydin Zahedivash, Alexander Tarlochan Singh Sandhu, Fatima Rodriguez, Daniel McDuff, Elgar Fleisch, Oliver Aalami, Filipe Barata, Paul Schmiedmayer

机构 * Stanford Mussallem Center for Biodesign(斯坦福 Mussallem 生物设计中心) Centre for Digital Health Interventions(数字健康干预中心) Agentic Systems Lab(代理系统实验室) National University of Singapore(新加坡国立大学) Microsoft(微软) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Google Research(谷歌研究) Stanford University(斯坦福大学) Amazon(亚马逊) Division of Cardiovascular Medicine(心血管医学部) Division of Cardiology(心内科部) Pediatric Cardiology(儿童心内科) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 OpenTSLM通过整合时间序列作为原生模态,提升对多变量医学文本和时间序列数据的推理能力,其模型在多个任务中均优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 11 篇

2602.11858 2026-02-17 cs.CV cs.AI cs.CL cs.LG 85%

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

无需缩放:用于细粒度多模态感知的区域到图像蒸馏

Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong, Yutong Cai, Siyuan Li, Huijia Zhu, Weiqiang Wang, Linghe Kong, Yue Wang, Zhuosheng Zhang, Weiran Huang

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Zhongguancun Academy(中关村学院) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出区域到图像蒸馏方法,通过训练时间内化代理缩放能力,提升细粒度多模态感知性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08417 2026-02-17 cs.RO 78%

ORACLE-Grasp: Zero-Shot Affordance-Aligned Robotic Grasping using Large Multimodal Models

ORACLE-Grasp: 基于大多模态模型的零样本 affordance 对齐机器人抓取

Avihai Giuili, Rotem Atari, Avishai Sintov

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 ORACLE-Grasp 利用大多模态模型实现零样本抓取,通过语义对齐提升抓取准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13653 2026-02-17 cs.AI cs.CL cs.CV cs.HC 75%

Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization

通过代理-Q估计和分步策略优化构建自主GUI导航

Yibo Wang, Guangda Huzhang, Yuwei Hu, Yu Xia, Shiyin Lu, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Lijun Zhang

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Ovis Team, Alibaba Group(阿里团队,阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出基于代理-Q估计和分步策略优化的GUI自主导航框架,通过强化学习提升GUI交互能力,实验证明其在导航和基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14979 2026-02-17 cs.RO 67%

RynnBrain: Open Embodied Foundation Models

RynnBrain: 开放式具身基础模型

Ronghao Dang, Jiayan Guo, Bohan Hou, Sicong Leng, Kehan Li, Xin Li, Jiangpin Liu, Yunxuan Mao, Zhikai Wang, Yuqian Yuan, Minghao Zhu, Xiao Lin, Yang Bai, Qian Jiang, Yaxi Zhao, Minghua Zeng, Junlong Gao, Yuming Jiang, Jun Cen, Siteng Huang, Liuyi Wang, Wenqiao Zhang, Chengju Liu, Jianfei Yang, Shijian Lu, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract)

AI总结 RynnBrain提出了一种开放的具身基础模型,通过统一框架强化感知、推理、规划等能力,显著优于现有模型,并能高效适应多种具身任务。

Comments Homepage: https://alibaba-damo-academy.github.io/RynnBrain.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14234 2026-02-17 cs.AI cs.CL 62%

REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents

REDSearcher: 一种可扩展且成本效益高的长周期搜索智能体框架

Zheng Chu, Xiao Wang, Jack Hong, Huiming Fan, Yuqi Huang, Yue Yang, Guohai Xu, Chenxiao Zhao, Cheng Xiang, Shengchao Hu, Dongdong Kuang, Ming Liu, Bing Qin, Xing Yu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 REDSearcher提出一种统一框架,通过任务合成、中期训练和后期训练优化长周期搜索智能体,实现高效且低成本的搜索任务解决。

Comments https://redsearchagent.github.io/index/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14093 2026-02-17 cs.AI cs.LG 57%

GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

GUI-GENESIS: 自动合成具有可验证奖励的高效环境以实现GUI代理训练

Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Xin Chen, Ang Li, Gang Cao, Gong Zhi, Hao Yu, Linyi Li, Wei Yang, Tao Xie

机构 * Key Lab of HCST (PKU), MOE SCS, Peking University, Beijing, China Tencent Inc., Shenzheng, China Hong Kong University of Science Department of Computer Science, University of Texas at Dallas, Richardson, USA School of Computing Science, Simon Fraser University, Burnaby, BC, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 GUI-GENESIS通过自动合成高效GUI训练环境并提供可验证奖励,显著提升了GUI代理的训练效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14048 2026-02-17 cs.RO cs.CV cs.GR 57%

ProAct: A Dual-System Framework for Proactive Embodied Social Agents

ProAct:一种双系统框架用于主动具身社交代理

Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng, Biao Jiang, Libin Liu

机构 * Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 ProAct通过双系统框架实现主动具身社交代理,结合低延迟行为系统与慢速认知系统,提升交互的主动性和社会参与度。

Comments Project Page: https://proactrobot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14003 2026-02-17 cs.AI 57%

Prompt-Driven Low-Altitude Edge Intelligence: Modular Agents and Generative Reasoning

基于提示的低空边缘智能:模块化代理与生成推理

Jiahao You, Ziye Jia, Chao Dong, Qihui Wu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出P2AECF框架,通过模块化代理和生成推理实现灵活、高效和适应的低空边缘智能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10080 2026-02-17 cs.CV 57%

BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals

BEVTraj: 无地图端到端鸟瞰图轨迹预测方法,采用可变形注意力和稀疏目标提案

Minsang Kong, Myeongjun Kim, Sang Gu Kang, Hejiu Lu, Yupeng Zhong, Sang Hun Lee

机构 * Department of Automobile and IT Convergence, Kookmin University(汽车与IT融合系,韩国釜山大学) Department of Automotive Engineering, Kookmin University(汽车工程系,韩国釜山大学) Graduate School of Automobile and Mobility, Kookmin University(汽车与移动研究生院,韩国釜山大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 BEVTraj通过可变形注意力和稀疏目标提案实现无地图端到端鸟瞰图轨迹预测,提升自动驾驶的鲁棒性和灵活性。

Comments Submitted to IEEE Transactions on Intelligent Transportation Systems (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13404 2026-02-17 astro-ph.EP physics.pop-ph 50%

The Interplanetary Habitable Zone

星际宜居带

Caleb Scharf

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种评估星际宜居带的多模式指标,并通过基于代理的模型探讨了星际生命扩散与资源利用之间的平衡。

Comments 38 pages, 14 color figures, submitted to The Astrobiology Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10285 2026-02-17 cs.RO 50%

Adaptive Time Step Flow Matching for Autonomous Driving Motion Planning

自适应时间步长流匹配用于自动驾驶运动规划

Ananya Trivedi, Anjian Li, Mohamed Elnoor, Yusuf Umut Ciftci, Avinash Singh, Jovin D'sa, Sangjae Bae, David Isele, Taskin Padir, Faizan M. Tariq

机构 * HRI Honda Research Institute(本田研究院) Northeastern University(东北大学) Princeton University(普林斯顿大学) University of Maryland(马里兰大学) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于条件流匹配的自适应时间步长框架,用于实时自动驾驶轨迹规划,通过在线调整推理步骤数和轨迹后处理提升性能。

Comments Accepted to Intelligent Vehicles Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 10 篇

2602.13715 2026-02-17 cs.IR 89%

DMESR: Dual-view MLLM-based Enhancing Framework for Multimodal Sequential Recommendation

DMESR: 基于双视角的多模态序列推荐增强框架

Mingyao Huang, Qidong Liu, Wenxuan Yang, Moranxin Wang, Yuqi Sun, Haiping Zhu, Feng Tian, Yan Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);cross-modal(abstract)

AI总结 DMESR通过双视角机制解决多模态序列推荐中的语义对齐和细粒度语义丢失问题,提升推荐效果。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13606 2026-02-17 cs.NI cs.AI cs.ET cs.LG 85%

Multi-Modal Sensing and Fusion in mmWave Beamforming for Connected Vehicles: A Transformer Based Framework

毫米波波束成形中多模态感知与融合:基于变换器的框架

Muhammad Baqer Mollah, Honggang Wang, Mohammad Ataul Karim, Hua Fang

机构 * Department of Electrical and Computer Engineering, University of Massachusetts Dartmouth(电子与计算机工程系,马萨诸塞大学达特茅斯分校) Department of Graduate Computer Science and Engineering, Katz School of Science and Health, Yeshiva University(研究生计算机科学与工程系,耶鲁大学科学与健康学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于变换器的多模态感知与融合框架,用于毫米波波束成形,以减少波束训练开销并提高连接车辆的通信效率。

Comments 13 Pages. arXiv admin note: text overlap with arXiv:2509.11112

Journal ref IEEE Transactions on Vehicular Technology, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13857 2026-02-17 cs.LG eess.SP 82%

sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals

sleep2vec:用于异构夜间生物信号的统一跨模态对齐

Weixuan Yuan, Zengrui Jin, Yichen Wang, Donglin Xie, Ziyi Ye, Chao Zhang, Xuesong Chen

机构 * Five Seasons Medical(五 Seasons 医疗) Tsinghua University(清华大学) Beijing Key Laboratory for Sleep Breathing Disorder(北京睡眠呼吸障碍重点实验室) Peking University(北京大学) Fudan University(复旦大学) Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

AI总结 sleep2vec通过统一跨模态对齐和原理化缩放,实现了对异构夜间生物信号的高效、通用建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14518 2026-02-17 cs.AI 79%

Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning

多模态长链推理中的知识冲突诊断

Jing Tang, Kun Wang, Haolang Lu, Hongjin Chen, KaiTao Chen, Zhongxiang Sun, Qiankun Li, Lingjuan Lyu, Guoshun Nan, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究提出了一种统一的知识冲突概念,揭示了多模态长链推理中不同冲突类型的特征和处理机制,为诊断和控制推理失败提供了原理性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10685 2026-02-17 cs.CV 79%

GaussianFormer3D: Multi-Modal Gaussian-based Semantic Occupancy Prediction with 3D Deformable Attention

GaussianFormer3D: 多模态基于高斯的语义占用预测与3D可变形注意力

Lingjun Zhao, Sizhe Wei, James Hays, Lu Gan

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 GaussianFormer3D通过3D可变形注意力机制,结合激光雷达与摄像头数据,实现高效且精确的语义占用预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13315 2026-02-17 cs.CV cs.AI 73%

IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs

IDPruner: 在视觉令牌修剪中协调重要性与多样性

Yifan Tan, Yifu Sun, Shirui Huang, Hong Liu, Guanghua Yu, Jianchen Zhu, Yangdong Deng

机构 * School of Software, Tsinghua University(清华大学软件学院) Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 IDPruner通过协调重要性与多样性,提出了一种高效的视觉令牌修剪方法,实现最优平衡并提升多模态大语言模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏