arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-20 至 2025-10-20 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2510.15261 2025-10-20 cs.AI 79%

AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory

Jitesh Jain, Shubham Maheshwari, Ning Yu, Wen-mei Hwu, Humphrey Shi

机构 * Adobe(Adobe公司) Netflix Eyeline Studios(Netflix Eyeline工作室) UIUC(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments LAW 2025 Workshop at NeurIPS 2025. Work done from late 2023 to early 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15767 2025-10-20 cs.SE 78%

EASELAN: An Open-Source Framework for Multimodal Biosignal Annotation and Data Management

Rathi Adarshi Rammohan, Moritz Meier, Dennis Küster, Tanja Schultz

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02444 2025-10-20 cs.AI cs.CL cs.CV cs.HC 75%

AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent

Jingru Fan, Yufan Dang, Jingyao Wu, Huatao Li, Runde Yang, Xiyuan Yang, Yuheng Wang, Chen Qian

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project at https://github.com/OpenBMB/AppCopilot

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11717 2025-10-20 cs.LG cs.AI cs.CL cs.CV 75%

WebInject: Prompt Injection Attack to Web Agents

Xilong Wang, John Bloch, Zedian Shao, Yuepeng Hu, Shuyan Zhou, Neil Zhenqiang Gong

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Appeared in EMNLP 2025 main conference. To better understand prompt injection attacks, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01055 2025-10-20 cs.AI cs.CL cs.CV 67%

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Dongfu Jiang, Yi Lu, Zhuofeng Li, Zhiheng Lyu, Ping Nie, Haozhe Wang, Alex Su, Hui Chen, Kai Zou, Chao Du, Tianyu Pang, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Sea AI Lab(Sea AI 实验室) University of Toronto(多伦多大学) Shanghai University(上海大学) HKUST(香港科技大学) M-A-P National University of Singapore(新加坡国立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 32 pages, 5 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17462 2025-10-20 cs.RO cs.AI cs.CV 62%

General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting

Bernard Lange, Anil Yildiz, Mansur Arief, Shehryar Khattak, Mykel Kochenderfer, Georgios Georgakis

机构 * Stanford University(斯坦福大学) Jet Propulsion Laboratory(喷气推进实验室) California Institute of Technology(加州理工学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏