arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2790 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2790 篇

2404.18083 2024-06-21 cs.RO cs.AI cs.CV 81%

Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching

Zhiwei Huang, Yikang Zhang, Qijun Chen, Rui Fan

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments accepted to IEEE Trans. on Intelligent Vehicles (T-IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06722 2024-06-12 cs.CV cs.CL cs.RO 81%

EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, Xihui Liu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project released at: https://github.com/ChenYi99/EgoPlan

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13919 2024-06-10 cs.CL cs.AI 81%

WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2024 (main). Code and data is released at https://github.com/MinorJerry/WebVoyager

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11436 2024-06-10 cs.CL cs.AI cs.HC 81%

You Only Look at Screens: Multimodal Chain-of-Action Agents

Zhuosheng Zhang, Aston Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13034 2024-06-07 cs.CL cs.AI cs.HC 81%

Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality

Jiahuan Pei, Irene Viola, Haochen Huang, Junxiao Wang, Moonisa Ahsan, Fanghua Ye, Jiang Yiming, Yao Sai, Di Wang, Zhumin Chen, Pengjie Ren, Pablo Cesar

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19267 2024-05-24 cs.CL cs.AI 81%

MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical Needs

Xianhao Yu, Jiaqi Fu, Renjia Deng, Wenjuan Han

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Project website: https://github.com/cocacola-lab/MineLand

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12222 2024-05-22 eess.IV cs.AI cs.CV 81%

Influence based explainability of brain tumors segmentation in multimodal Magnetic Resonance Imaging

Tommaso Torda, Andrea Ciardiello, Simona Gargiulo, Greta Grillo, Simone Scardapane, Cecilia Voena, Stefano Giagu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03627 2024-04-29 cs.CL cs.AI 81%

Multimodal Large Language Models to Support Real-World Fact-Checking

Jiahui Geng, Yova Kementchedjhieva, Preslav Nakov, Iryna Gurevych

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11330 2024-04-24 cs.CL cs.AI cs.HC cs.LG 81%

Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback

Dong Won Lee, Hae Won Park, Yoon Kim, Cynthia Breazeal, Louis-Philippe Morency

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06682 2024-03-12 cs.CL cs.CV cs.CY 81%

Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach

Siyu Duan, Jun Wang, Qi Su

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accept by Lrec-Coling 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11029 2024-03-04 cs.CL cs.AI 81%

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02330 2024-02-23 cs.CV cs.CL 81%

LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Yichen Zhu, Minjie Zhu, Ning Liu, Zhicai Ou, Xiaofeng Mou, Jian Tang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments The datasets were incomplete as they did not include all the necessary copyrights

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03610 2024-02-07 cs.LG cs.AI cs.CL 81%

RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents

Tomoyuki Kagaya, Thong Jing Yuan, Yuxuan Lou, Jayashree Karlekar, Sugiri Pranata, Akira Kinose, Koki Oguri, Felix Wick, Yang You

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14859 2023-12-22 cs.CV cs.CL 81%

3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction

Mehdi Fatan, Emanuele Mincato, Dimitra Pintzou, Mariella Dimiccoli

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07886 2023-12-14 cs.AI cs.CL 81%

Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI

Kai Huang, Boyuan Yang, Wei Gao

专题命中 多模态Agent :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07562 2023-11-14 cs.CV cs.AI 81%

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, Zicheng Liu, Lijuan Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04067 2023-11-08 cs.LG cs.AI cs.CV 81%

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10790 2023-10-26 cs.LG cs.AI cs.CV cs.RO 81%

Guide Your Agent with Adaptive Multimodal Rewards

Changyeon Kim, Younggyo Seo, Hao Liu, Lisa Lee, Jinwoo Shin, Honglak Lee, Kimin Lee

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS 2023. Project webpage: https://sites.google.com/view/2023arp

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13850 2023-07-27 cs.LG cs.AI cs.CV cs.RO 81%

MAEA: Multimodal Attribution for Embodied AI

Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, Yonatan Bisk

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06485 2023-05-12 cs.RO cs.AI cs.CL cs.HC 81%

Multimodal Contextualized Plan Prediction for Embodied Task Completion

Mert İnan, Aishwarya Padmakumar, Spandana Gella, Patrick Lange, Dilek Hakkani-Tur

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments NILLI at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09151 2020-08-24 cs.CL cs.MM 81%

Multi-modal Cooking Workflow Construction for Food Recipes

Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang, Tat-Seng Chua

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.MM

Comments This manuscript has been accepted at ACM MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01385 2019-02-05 cs.LG cs.AI cs.CL cs.RO stat.ML 81%

Embodied Multimodal Multitask Learning

Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh, Dhruv Batra

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments See https://devendrachaplot.github.io/projects/EMML for demo videos

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11954 2018-11-22 cs.CL cs.AI 81%

A Knowledge-Grounded Multimodal Search-Based Conversational Agent

Shubham Agarwal, Ondrej Dusek, Ioannis Konstas, Verena Rieser

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, pages 59-66, Brussels, Belgium, October 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16774 2025-10-21 cs.LG cs.AI 80%

Learning to play: A Multimodal Agent for 3D Game-Play

Yuguang Yue, Irakli Salia, Samuel Hunt, Christopher Green, Wenzhe Shi, Jonathan J Hunt

机构 * Player2

专题命中 多模态Agent :multimodal(title);multi-modal(abstract,comments);分类 cs.AI

Comments International Conference on Computer Vision Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.07025 2020-09-16 cs.CV 80%

FairCVtest Demo: Understanding Bias in Multimodal Learning with a Testbed in Fair Automatic Recruitment

Alejandro Peña, Ignacio Serna, Aythami Morales, Julian Fierrez

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments ACM Intl. Conf. on Multimodal Interaction (ICMI). arXiv admin note: substantial text overlap with arXiv:2004.07173

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19436 2026-08-12 cs.CV cs.AI cs.LG cs.MM 版本更新 80%

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

VDC-Agent:当视频详细描述器通过代理自我反思而自我进化

Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong

机构 * Xi’an Jiaotong University(西安交通大学) Kuaishou Technology(快手科技) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 VDC-Agent通过自我反思机制实现视频详细描述的自我进化,利用自动生成的(描述,评分)对提升描述准确性与评分表现。

Comments Accepted to ECCV 2026. Project Page: https://vdcagent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26775 2026-08-04 cs.LG cs.AI cs.CL cs.CV 80%

Learning to Select Visual In-Context Demonstrations

学习选择视觉上下文示例

Eugene Lee, Yu-Chi Lin, Jiajie Diao

机构 * University of Cincinnati(辛辛那提大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出LSD方法,通过强化学习构建最优演示集,提升多模态大语言模型在视觉回归任务中的表现,揭示了学习选择在视觉上下文学习中的必要性。

Comments 21 pages, 12 figure, accepted to Computer Vision and Pattern Recognition Conference (CVPR) 2026 Findings Track

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9455-9465) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27926 2026-06-29 cs.AI cs.CL cs.CV 新提交 80%

Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing

可验证几何问题求解:求解器驱动的自动形式化与定理提出

Can Li, Ting Zhang, Junbo Zhao, Hua Huang

机构 * Beijing Normal University(北京师范大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出SD-GPS框架,通过求解器驱动的自动形式化和可验证定理提出,解决几何问题求解中神经符号方法的瓶颈,在Geometry3K和PGPS9K上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16491 2026-06-16 cs.RO 新提交 80%

HATS: A Human-Agent Teleoperation System for Multi-Arm Data Collection

HATS:用于多臂数据收集的人-智能体遥操作系统

Zesen Lin, Jian-Jian Jiang, Haoming Cen, Xiao-Ming Wu, Dandan Zhang, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Nanyang Technological University(南洋理工大学) Imperial College London(帝国理工学院)

专题命中 多模态Agent :MLLM(summary_cn,abstract)

AI总结 提出HATS系统,由单操作员借助MLLM智能体控制两主臂和两辅助臂,实现高效多臂数据收集,性能媲美双人专家团队。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20237 2026-08-21 cs.AI 新提交 79%

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

面向多模态大语言模型的符合规则的视觉空间规划

Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所) Yinwang Intelligent Technology Co., Ltd(银湾智能科技有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对多模态大语言模型的规则遵循空间规划问题,构建了RuleMaze基准,提出语言-逻辑-函数混合方法和解耦多模态规划(DMP),提升了规则遵循度与规划成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏