arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-09 至 2025-12-09 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 10 篇

2512.06396 2025-12-09 cs.CR cs.AI 83%

AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity

AgenticCyber: 一种基于生成式AI的多智能体系统,用于多模态威胁检测与自适应响应在网络安全中

Shovan Roy

机构 * TNTech

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 AgenticCyber通过生成式AI驱动的多智能体系统,实现多模态威胁检测与自适应响应,提升网络安全性能和态势感知能力。

Comments 6 pages for IEEE conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07351 2025-12-09 cs.CV cs.AI cs.SD 81%

DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection

DeepAgent: 一种双流多智能体融合用于鲁棒多模态深度伪造检测

Sayeem Been Zaman, Wasimul Karim, Arefin Ittesafun Abian, Reem E. Mohamed, Md Rafiqul Islam, Asif Karim, Sami Azam

机构 * Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室) Department of Computer Science and Engineering(计算机科学与工程系) University of Scholars(学者大学) Faculty of Science and Information Technology(科学与信息技术学院) Faculty of Science and Technology(科学与技术学院)

专题命中 多模态Agent :multimodal(title);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 DeepAgent通过双流多智能体融合方法,提升多模态深度伪造检测的鲁棒性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02395 2025-12-09 cs.CV 79%

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Skywork-R1V4:通过图像与深度研究交织思考实现代理多模态智能

Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang, Yichen Wei, Shukang Yin, Jiangbo Pei, Wei Shen, Peng Xia, Yi Peng, Tianyidan Xie, Eric Li, Yang Liu, Xuchen Song, Yahui Zhou

机构 * Skywork AI

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 Skywork-R1V4通过交织推理实现多模态代理智能,仅用监督学习在少数据上训练,超越现有模型在多个基准测试中的表现。

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07132 2025-12-09 cs.CL cs.AI cs.CV 78%

DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning

利用多智能体分歧进行多模态推理中的工具招募

Nithin Sivakumaran, Justin Chih-Yao Chen, David Wan, Yue Zhang, Jaehong Yoon, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.CL、cs.AI

AI总结 DART通过多智能体分歧识别有用视觉工具,提升多模态推理中的工具调用效果。

Comments Code: https://github.com/nsivaku/dart

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05731 2025-12-09 cs.AI cs.CL 62%

InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization

InfiGUI-G1: 通过自适应探索策略优化推进GUI接地

Yuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li, Congkai Xie, Jiasheng Wang, Xavier Hu, Xiaotian Han, Jianbo Yuan, Xinyao Wang, Shengyu Zhang, Hongxia Yang, Fei Wu

机构 * Amazon(亚马逊)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 InfiGUI-G1通过自适应探索策略优化改进GUI接地,实现显著的语义对齐和性能提升。

Comments Accepted to AAAI 2026 (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02292 2025-12-09 cs.AI cs.LG 57%

FinWorld: An All-in-One Open-Source Platform for End-to-End Financial AI Research and Deployment

FinWorld: 一个全方位的开源平台,用于端到端的金融AI研究与部署

Wentao Zhang, Yilei Zhao, Chuqiao Zong, Xinrun Wang, Bo An

机构 * Nanyang Technological University Skywork AI(南洋理工大学Skywork AI) Nanyang Technological University(南洋理工大学) Singapore Management University(新加坡管理学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 FinWorld是一个全方位的开源平台,提供端到端的金融AI研究与部署支持,通过整合异构金融数据和强化学习技术,提升可重复性和部署效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06373 2025-12-09 cs.CV 57%

VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning

VG-Refiner: 通过代理强化学习实现工具精细化的指称 grounded 推理

Yuji Wang, Wenlong Liu, Jingxuan Niu, Haoji Zhang, Yansong Tang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) International Digital Economy Academy (IDEA)(国际数字经济学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 VG-Refiner通过代理强化学习实现工具精细化的指称 grounded 推理,引入两阶段思考-重新思考机制和细化奖励,提升指称和推理接地任务的准确性和修正能力。

Comments The project page is [this url](https://github.com/VoyageWang/VG-Refiner)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07452 2025-12-09 cs.IR 50%

From Show Programmes to Data: Designing a Workflow to Make Performing Arts Ephemera Accessible Through Language Models

从节目到数据:设计一种工作流,通过语言模型使表演艺术的临时性资料可访问

Clarisse Bardiot, Pierre-Carl Langlais, Bernard Jacquemin, Jacob Hart, Antonios Lagarias, Nicolas Foucault, Aurélie Lemaître-Legargeant, Jeanne Fras

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出利用多模态语言模型和本体推理模型,将戏剧节目转化为结构化数据,以提升文化遗产资料的可访问性和分析能力。

Comments 19 pages, 8 figures, 5 tables, 17 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07032 2025-12-09 cs.RO 50%

A Hetero-Associative Sequential Memory Model Utilizing Neuromorphic Signals: Validated on a Mobile Manipulator

一种利用神经形态信号的异关联序列记忆模型:在移动机械臂上的验证

Runcong Wang, Fengyi Wang, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出一种利用神经形态信号的异关联序列记忆模型,用于移动机械臂的触觉与动作决策,实现低计算成本的伪柔顺控制与多关节抓取序列检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06558 2025-12-09 cs.RO 50%

Embodied Referring Expression Comprehension in Human-Robot Interaction

具身指称表达理解在人机交互中的应用

Md Mofijul Islam, Alexi Gladstone, Sujan Sarker, Ganesh Nanduru, Md Fahim, Keyan Du, Aman Chadha, Tariq Iqbal

机构 * University of Virginia(弗吉尼亚大学) Stanford University(斯坦福大学) University of Dhaka(达卡大学) Amazon GenAI(亚马逊生成人工智能)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出Refer360数据集和MuRes模块,用于提升机器人在人机交互中对具身指称表达的理解能力。

Comments 14 pages, 7 figures, accepted at the ACM/IEEE International Conference on Human-Robot Interaction (HRI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏