arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-05 至 2026-02-05 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 8 篇

2602.04486 2026-02-05 cs.CL 83%

Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition

超越单模态捷径:MLLMs作为跨模态推理器用于 grounded 命名实体识别

Jinlong Ma, Yu Zhang, Xuefeng Bai, Kehai Chen, Yuwei Wang, Zeming Liu, Jun Yu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Beijing University of Aeronautics and Astronautics(北京航空航天大学)

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

AI总结 本文提出MCR方法,通过多风格推理模式注入和约束引导的可验证优化,解决MLLMs在跨模态推理中的模态偏见问题,提升 grounded 命名实体识别的性能。

Comments GMNER

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07825 2026-02-05 cs.CV cs.AI cs.LG 81%

Deep Multimodal Learning with Missing Modality: A Survey

缺失模态下的深度多模态学习:综述

Renjie Wu, Hu Wang, Hsiang-Ting Chen, Gustavo Carneiro

机构 * The Australian National University(澳大利亚国立大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Adelaide University(阿德莱德大学) The University of Surrey(萨里大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文综述了缺失模态下的多模态学习方法,分析了其动机、技术细节、应用及挑战,为该领域的发展提供了全面的视角。

Comments Accepted by TMLR (Transactions on Machine Learning Research)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14265 2026-02-05 cs.LG 78%

Unified Multimodal Vessel Trajectory Prediction with Explainable Navigation Intention

具有可解释导航意图的统一多模态船舶轨迹预测

Rui Zhang, Chao Li, Kezhong Liu, Chen Wang, Bolong Zheng, Hongbo Jiang

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Navigation, Wuhan University of Technology(武汉理工大学航海学院) Hubei Key Laboratory of Internet of Intelligence, School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院智能互联网关键实验室) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出了一种整合可解释导航意图的统一多模态船舶轨迹预测框架,通过构建持续性意图树和动态瞬时意图模型,提升预测的准确性和可解释性。

Journal ref IEEE Transactions on Intelligent Transportation Systems, vol. 27, no. 1, pp. 258-269, Jan. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03949 2026-02-05 cs.IT cs.AI cs.LG math.IT 57%

Semantic Rate Distortion and Posterior Design: Compute Constraints, Multimodality, and Strategic Inference

语义速率与后验设计:计算限制、多模态与战略推断

Emrah Akyol

机构 * Electrical and Computer Engineering Department, Binghamton University(宾夕法尼亚大学布林茅尔分校电子与计算机工程系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了在计算限制下的语义压缩问题,通过后验设计和多模态观测优化,提高了语义准确性和模型效率。

Comments submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03908 2026-02-05 cs.RO cs.CV 57%

Beyond the Vehicle: Cooperative Localization by Fusing Point Clouds for GPS-Challenged Urban Scenarios

超越车辆:通过融合点云进行协作定位以应对GPS挑战的都市场景

Kuo-Yi Chao, Ralph Rasshofer, Alois Christian Knoll

机构 * Technical University of Munich(慕尼黑技术大学) BMW Group(宝马集团)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种融合点云的协作定位方法,通过多传感器和多模态数据提升GPS不可靠城市环境中的定位精度和鲁棒性。

Comments 8 pages, 2 figures, Driving the Future Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15605 2026-02-05 cs.CL cs.SI 57%

ToxiTwitch: Toward Emote-Aware Hybrid Moderation for Live Streaming Platforms

ToxiTwitch:迈向面向表情的混合审核方法

Baktash Ansari, Elias Martin, Afra Mashhadi

机构 * University of Washington(华盛顿大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 ToxiTwitch通过结合LLM生成的文本和表情嵌入与传统机器学习分类器,提高Twitch直播平台对有毒行为的检测准确率。

Comments Exploratory study; prior versions submitted to peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04736 2026-02-05 stat.ML cs.LG 50%

Conditional Counterfactual Mean Embeddings: Doubly Robust Estimation and Learning Rates

条件反事实均值嵌入:双重稳健估计与学习速率

Thatchanon Anancharoenkij, Donlapark Ponnoprat

机构 * Chiang Mai University(恰良迈大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出条件反事实均值嵌入框架,结合双重稳健估计与学习速率分析,用于异质处理效应的估计与学习。

Comments Code is available at https://github.com/donlap/Conditional-Counterfactual-Mean-Embeddings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19880 2026-02-05 econ.GN cs.GT math.OC q-fin.EC 50%

Mobility-as-a-service (MaaS) system as a multi-leader-multi-follower game: A single-level variational inequality (VI) formulation

作为多领导者-多追随者博弈的移动即服务(MaaS)系统:一种单层变分不等式(VI) formulations

Rui Yao, Xinyu Ma, Kenan Zhang

专题命中 其他多模态 :multi-modal(abstract)

AI总结 本文提出了一种单层变分不等式 formulations,用于建模多领导者-多追随者博弈的MaaS系统,通过引入虚拟交通运营商实现并行求解,验证了模型的可扩展性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏