arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2109.07951 2021-09-17 cs.CV 80%

Overview of Tencent Multi-modal Ads Video Understanding Challenge

Zhenzhi Wang, Liyu Wu, Zhimin Li, Jiangfeng Xiong, Qinglin Lu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8-page extended version of our challenge paper in ACM MM 2021. It presents the overview of grand challenge "Multi-modal Ads Video Understanding" in ACM MM 2021. Our grand challenge is also the Tencent Advertising Algorithm Competition (TAAC) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13570 2019-11-25 cs.LG cs.AI cs.NE stat.ML 80%

Factorized Inference in Deep Markov Models for Incomplete Multimodal Time Series

Tan Zhi-Xuan, Harold Soh, Desmond C. Ong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 4 figures, accepted to AAAI 2020, code available at: https://github.com/ztangent/multimodal-dmm

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02908 2019-03-14 cs.RO cs.CV 80%

Multi-Modal Obstacle Detection in Unstructured Environments with Conditional Random Fields

Mikkel Kragh, James Underwood

专题命中 视频多模态 :multi-modal(title);multimodal(abstract,comments);分类 cs.CV

Comments This is the accepted version of the following article: Kragh M, Underwood J. Multimodal obstacle detection in unstructured environments with conditional random fields. J Field Robotics. 2019, 1-20., which has been published in final form at https://doi.org/10.1002/rob.21866

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00721 2018-05-03 cs.CV 80%

Joint Surgical Gesture and Task Classification with Multi-Task and Multimodal Learning

Duygu Sarikaya, Khurshid A. Guru, Jason J. Corso

专题命中 视频多模态 :multimodal(title,comments);multi-modal(abstract);分类 cs.CV

Comments Keywords Robot-Assisted Surgery, Surgical Gesture Classification, Multi-task Learning, Multimodal Learning, Long Short-term Recurrent Neural Networks, Convolutional Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.05212 2016-05-18 cs.LG cs.CV 80%

Multimodal Sparse Coding for Event Detection

Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim, Miriam Cha, H. T. Kung

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Multimodal Machine Learning Workshop at NIPS 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13210 2026-08-14 cs.CV cs.AI cs.MM 新提交 80%

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU:用于理解日语超长视频中叙事演变与文化细微差别的基准

Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma

机构 * The University of Tokyo(东京大学) Kyushu University(九州大学) Macau University of Science and Technology(澳门科技大学) Infinimind Japan Inc.(Infinimind日本公司) University of Alberta(阿尔伯塔大学)

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI、cs.MM

AI总结 该研究推出NARU基准,涵盖155个146.8小时日语视频的1481个问题,评估模型在长程叙事整合与文化推理上的局限,为MLLM开发提供测试平台。

Comments Yuheng Huang and Jianlang Chen contributed equally to this work. More details available on the project's website https://ma-labo.github.io/naru/ and https://infinimind.io/en/company/news/2026/narubench-release

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12290 2026-08-13 cs.CV cs.AI cs.MM 新提交 80%

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

超越试错:面向图像到视频一致性的智能体优化

Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson

机构 * Google Cloud(谷歌云) Google DeepMind(谷歌DeepMind)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 针对图像到视频模型试错效率低的问题,提出Agentic Self-Improvement框架,通过两阶段优化提升视频与文本一致性,生成视频胜率达69%,为视频生成模型提供实用可控的优化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11534 2026-08-11 cs.CV cs.CL cs.MM 版本更新 80%

HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding

HFS: 为高效视频推理的全局查询感知帧选择

Yiqing Yang, Yun Li, Daiqing Qi, Lehan Yang, Tianlong Wang, Wenhao Zhang, Sheng Li, Kin-man Lam

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 HFS提出一种端到端可训练的帧选择框架,通过任务自适应方法提升视频推理效率。

Comments Accepted to the Main Track of ACM Multimedia 2026 (ACM MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04515 2026-08-06 cs.CV cs.AI cs.CL 新提交 80%

CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding

CARVE:用于高效3D医学体积理解的视觉证据跨切片各向异性重分配

Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 针对3D医学体积理解中切片式MLLM的视觉令牌冗余问题,提出无需训练的CARVE框架,通过跨切片各向异性重分配压缩80%令牌,在AMOS-MM等基准上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19378 2025-08-05 cs.CV cs.AI cs.CL cs.LG 80%

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * Information Retrieval Group(信息检索组) AI4BioMed Lab(AI4BioMed实验室) School of Computing Science(计算科学学院) University of Glasgow(格拉斯哥大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 30 pages, 5 figures, Adding Appendix

Journal ref Association for Computational Linguistics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11701 2022-11-22 cs.CV cs.AI cs.CL 80%

Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention

Zineng Tang, Jaemin Cho, Jie Lei, Mohit Bansal

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments WACV 2023 (first two authors contributed equally)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23503 2026-08-25 cs.CV 新提交 79%

Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search

用于基于文本的人员异常搜索的动作对齐检索与成对多模态重排序

Thanh-Khoi Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen, Minh-Triet Tran

机构 * University of Science, VNU-HCM(胡志明市国家大学科学大学) Vietnam National University, Ho Chi Minh City(胡志明市越南国家大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 针对基于文本的人员异常搜索中现有方法的局限,提出ActPair三阶段由粗到细框架,结合动作对齐检索与成对多模态重排序,在PAB测试集取得最优结果且可有效迁移至新数据集。

Comments Accepted to the AI City workshop @ ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22965 2026-08-25 cs.CV 新提交 79%

Simplified Cross-Modal Calibration for Heterogeneous Event-RGB Stereo Systems

面向异构事件-RGB立体系统的简化跨模态标定

Nico Hessenthaler, Adam T. Müller, Nicolaj C. Stache

机构 * Center for Machine Learning Heilbronn University of Applied Sciences(海尔布隆应用科学大学机器学习中心)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对异构事件-RGB立体系统的标定瓶颈,提出无运动跨模态标定框架,简化流程且降低重投影误差,在机器人眼到手标定中具实用性。

Comments Accepted to the 37th British Machine Vision Conference (BMVC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22869 2026-08-25 cs.RO cs.CV 新提交 79%

UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

UniMem:统一视觉-语言-动作模型的多模态记忆与控制

Lars Osterberg, Maggie Wang, Mac Schwager

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 UniMem 是统一多模态记忆与控制的视觉-语言-动作模型框架,通过事件分类器、关键帧编码器等技术,在模拟和硬件任务中性能优于基线,推理更快、训练流程更简单。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22832 2026-08-25 cs.AI 新提交 79%

Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku

让子弹飞:基于时间对齐生成式弹幕的多模态假新闻检测

Xiansheng Luo, Chaowei Zhang, Zewei Zhang, Yi Zhu, Jipeng Qiang

机构 * Yangzhou University(扬州大学) Auburn University(奥本大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出Genda生成式弹幕框架与DM-FEND检测模型,解决弹幕延迟导致的假新闻检测实时性问题,在FakeSV、FakeTT基准上性能优于现有最优方法,为多模态假新闻检测提供鲁棒方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14935 2026-08-25 cs.CV 版本更新 79%

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

VideoChat3:用于高效通用视频理解的全开放视频多模态语言模型

Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) Peking University(北京大学)

专题命中 视频多模态 :MLLM(title,abstract);分类 cs.CV

AI总结 研究针对视频理解开源模型局限,提出VideoChat3。通过I3D-ViT等提升效率,利用可扩展视频数据合成管道生成训练数据集提升泛化性,以4B参数在多基准测试中超越同等或更多参数的开源模型,实现泛化与计算效率平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02846 2026-08-21 cs.CV 79%

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

瞬间动作预见:多模态线索能替代视频到何种程度?

Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Universidad de Alicante(阿利坎特大学) Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会) University of Crete(克里特大学) Hellenic Mediterranean University(希腊地中海大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 AAG通过结合单帧RGB特征与深度线索及先前动作信息,实现了多模态单帧动作预见,能与视频聚合基线和先进方法在教学活动数据集上竞争。

Comments Accepted in WACV 2026 - Applications Track

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22854 2026-08-20 cs.CV 79%

CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting

Kornel Howil, Joanna Waczyńska, Piotr Borycki, Tadeusz Dziarmaga, Marcin Mazur, Przemysław Spurek

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14657 2026-08-18 cs.LG cs.CV 新提交 79%

LUNG-KGMM: Knowledge-Guided Multimodal Learning for Lung Cancer Incidence Prediction

LUNG-KGMM:用于肺癌发病预测的知识引导多模态学习

Chunlei Yang, Shuyan Li, Zhong Cao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出LUNG-KGMM知识引导多模态框架,整合多类数据与临床知识开展1-6年肺癌发病预测,经MIMIC及厦门队列验证,其性能优于现有方法且具备跨队列可迁移性。

Comments 22 pages, 4 figures, 7 tables, accepted by PRCV Oral

Journal ref The 9th Chinese Conference on Pattern Recognition and Computer Vision, PRCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21333 2026-08-18 cs.CV 版本更新 79%

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

MME-VideoOCR:评估多模态大语言模型在视频场景中基于OCR的能力

Yang Shi, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yi-Fan Zhang, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Wenting Liu, Zhuoran Zhang, Xinlong Chen, Bohan Zeng, Sihan Yang, Yushuo Guan, Zhang Zhang, Liang Wang, Haoxuan Li, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Wenjing Yang

机构 * Kling Team(Kling团队) PKU(北京大学) THU(清华大学) CASIA(中国科学院自动化研究所) CUHKSZ(香港科技大学) NTU(国立台湾大学) NJU(南京大学) XJTU(吉林大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文推出MME-VideoOCR视频OCR基准,评估18个最先进MLLMs的视频OCR能力,发现现有模型在需视频整体理解的任务中表现有限,最优模型准确率仅73.7%,明确了高分辨率输入等对视频OCR的重要性。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14157 2026-08-17 cs.AI cs.LG 新提交 79%

Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

消除时间性笔记冗余可提升医学多模态强化学习性能

Chenran Weng, Joo Seung Lee, Malini Mahendra, Anil Aswani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对机械通气决策的RL方法缺失笔记临床信息的问题,提出冗余感知多模态框架,通过两种策略去除笔记冗余,在ICU数据上显著提升了RL临床决策性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12398 2026-08-14 q-bio.NC cs.AI 新提交 79%

A Hierarchical Energy-Based Model for Multimodal Cognition

面向多模态认知的分层能量基模型

Subir Varma

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究提出IM-LEPP多模态能量基模型,扩展单模态LEPP整合视觉与语言,能解释注意现象、复现心理语言学发现,还与Transformer模型形成可证伪对比并提出实验预测。

Comments 48 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13092 2026-08-14 cs.CV 新提交 79%

Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification

Paths:面向RGB-Event视频行人重识别的感知提示的时空Transformer与分层多模态融合

Yakun Huo, Yingquan Wang, Yangyang Liu, Tianyu Yan, Yunzhi Zhuge, Pingping Zhang, Huchuan Lu

机构 * Dalian University of Technology(大连理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对现有RE-VReID方法时空建模解耦、多模态融合不足的问题,提出含MAB、PST、HMF模块的Paths框架,在三个基准上验证了有效性。

Comments Accepted by ACM MM2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12920 2026-08-14 cs.CV 新提交 79%

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

TennisVAR:一种基于击球证据的多模态大语言模型,用于网球视频中的战术推理

Yifan Mei, Qingling Shi, Changli Wu, Jiayuan Rao, Jiayi Ji, Liujuan Cao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对网球视频理解的感知与理解差距,提出回合级战术推理任务,构建专家标注基准TRACE,开发基于证据的多模态大语言模型TennisVAR,实现网球视频的战术推理。

Comments Project Page: https://whynotgit2025.github.io/TennisVAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17598 2026-08-13 cs.RO cs.CV 版本更新 79%

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

MuseVLA: 一种用于机器人操作的自适应多模态感知视觉-语言-动作模型

Xingyuming Liu, Ruichun Ma, Heyu Guo, Qixiu Li, Qingwen Yang, Lin Luo, Shiqi Jiang, Chenren Xu, Jiaolong Yang, Baining Guo

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) Microsoft Research Asia(微软亚洲研究院) Princeton University(普林斯顿大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 提出MuseVLA模型,通过将传感器作为按需工具集成,实现自适应多模态感知;设计传感器图像统一表示,并引入数据合成流水线,在灵巧手操作任务中平均成功率80.6%,显著优于RGB-only和多模态基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10368 2026-08-12 cs.HC cs.MM 新提交 79%

Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction

XR中的视觉到触觉增强:一种用于多模态交互感知接地的可穿戴手套

Faisal Mohd, Hamdi Elsaddik, Erhan Baturay Onural, Jihong Zhang, Fedwa Laamarti, Abdulmotaleb El Saddik

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

AI总结 该研究针对XR系统触觉感知利用不足的问题,提出一款可穿戴手套及视觉到触觉映射算法,经20人实验验证,其触觉增强可提升XR交互的真实感与沉浸感,为多模态XR系统提供感知增强层。

Comments 11 pages, 4 figures. Published in the Proceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), April 13, 2026, Barcelona, Spain

Journal ref Proceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), CEUR Workshop Proceedings, Vol. 4226, pp. 252-262, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08915 2026-08-11 cs.CL 新提交 79%

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

视频介导对话中不同伙伴可见性条件下的多模态信息性研究

Esam Ghaleb, Hugh Mee Wong, Kristina Kobrock

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究针对视频介导对话,构建基于语音、手势或两者的模型识别指称对象,发现手势可单独预测指称,融合模型在转录模型不确定时效果最佳,还揭示了伙伴可见性对互动的语用效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07512 2026-08-11 cs.HC cs.AI 新提交 79%

EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews

EMMR:异步视频面试中用于人格评估的情绪介导多模态推理

Dongsheng Hu, Tianyi Zhang, Chuang Liu, Yuan Zong Yong Li, Wenming Zheng, Xiu-xiu Zhan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对异步视频面试人格评估中以文本为中心的方法忽略非言语线索的问题,提出EMMR框架,在OPVA和AVI-6数据集上提升了MAE、MSE和PCC,证明整合情绪线索可优化评估效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06973 2026-08-10 cs.CV 新提交 79%

When One Modality Is Not Enough: Multimodal Sex and Life-Stage Classification of Red Deer from Aerial RGB-Thermal Video

当单模态不够时:基于航拍RGB-热红外视频的马鹿多模态性别与生命阶段分类

Hugo Markoff, Christoph Praschl, Ivan Ludoški, Sara Beery, Michael Ørsted, David C. Schedl

机构 * Aalborg University(奥尔堡大学) University of Applied Sciences Upper Austria(上奥地利应用科学大学) University of Novi Sad(诺威萨德大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究以马鹿为对象,融合航拍RGB与热红外视频的自监督DINOv3特征,构建多模态分类流程,在四次飞行中实现26只个体25只的正确分类,提升了种群性别与生命阶段分类的鲁棒性,可自动化获取兽群结构数据。

Comments Accepted at the ECCV 2026 Workshop on Computer Vision for Ecology (CV4Ecology), archival proceedings track. 17 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05864 2026-08-07 cs.AI 新提交 79%

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

视觉感知不等于决策:多模态大语言模型能成为有效的首席执行官吗?

Yuyang Dai, Xueqing Peng, Yuxia Wang, Preslav Nakov, Zhuohan Xie

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究构建C-SUITEBENCH多模态基准,评估9个前沿模型作为CEO的决策能力,发现多模态输入可提升证据推理但会损害受限资源分配,揭示多模态智能体的视觉感知与受限行动是可分离瓶颈。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏