arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-08-25 至 2026-08-25 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 18 篇

2608.23268 2026-08-25 cs.CV 新提交 85%

Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner

面向多模态智能体学习者的双粒度智能体记忆与Shapley上下文归因

Jieke Wang, Tiancheng Shen, Yibo Yang, Ming-Hsuan Yang

机构 * UC Merced(加州大学默塞德分校) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 针对多模态大模型推理不足的问题,提出双粒度智能体记忆框架DG-Mem,结合Shapley上下文归因,在多个数学多模态推理数据集上实现性能提升,且无需参数更新即可适配各类骨干模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29902 2026-08-25 cs.AI 版本更新 85%

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP-Bench: 向多模态大语言模型交错生成的代理工具规划迈进

Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang

机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴通义千问大模型应用团队) Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学)

专题命中 多模态Agent :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 针对多模态大语言模型交错生成中事实性与创造性难以统一的问题,提出ATP-Bench基准,包含7702个问题-答案对,评估代理工具规划能力,揭示模型在交错规划中的不足。

Comments Accepted at the European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22885 2026-08-25 cs.CV 新提交 84%

DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

DRAgent:用于指代表达分割的判别推理智能体

Yujie Qi, Luyan Zhang

机构 * School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院) Khoury College of Computer Sciences, Northeastern University(东北大学Khoury计算机科学学院)

专题命中 多模态Agent :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文针对指代表达分割中MLLM单次坐标预测导致的定位偏差问题,提出DRAgent判别推理框架,通过两阶段目标选择与LoRA微调提升性能,在多个基准数据集上表现具竞争力。

Comments 5 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19739 2026-08-25 cs.CV cs.AI cs.LG 版本更新 84%

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

面向多模态视觉问答的问题引导式证据获取

Alin-Ionut Popa

机构 * Amazon Inc.(亚马逊公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对多模态视觉问答中模型读取文档不可靠的问题,提出Q-Guide智能体引导感知,在两个数据集上优于基线方法,且增益来自精准感知引导而非复杂控制逻辑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23435 2026-08-25 cs.CV cs.AI 新提交 82%

Towards Comprehensive Basketball Understanding

迈向全面的篮球理解

Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, Weidi Xie

专题命中 多模态Agent :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 针对现有篮球理解基准仅单一评估能力的不足,构建多模态基准BasketballBench,并提出智能体BasketballSkills,实验显示其在整合多能力的篮球理解任务上优于现有MLLM。

Comments 26 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21868 2026-08-25 cs.AI 新提交 79%

HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews

HiMA-MDD:用于临床访谈中可解释多模态抑郁检测的分层多智能体框架

Ao Chen, Xiaojiang Peng

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本研究针对临床访谈多模态抑郁检测的分层评估需求,提出分层多智能体框架HiMA-MDD,其分三层智能体处理流程,在E-DAIC数据集上以Qwen2.5-72B-Instruct为骨干,性能优于现有最优方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17642 2026-08-25 cs.AI 版本更新 79%

FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness

FinAcumen: 通过自演化经验记忆实现的金融多模态推理

Pianran Guo, Pengcheng Zhou, Yucheng Jian, Shuhua Chen, Zhongliang Yang, Linna Zhou

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Queen Mary University of London(伦敦玛丽女王大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 提出FinAcumen框架,通过选择性经验记忆机制增强工具增强型多模态推理,在四个金融基准上持续提升冻结的8B视觉语言模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22639 2026-08-25 cs.HC 新提交 78%

Poetic Heritage for Culturally Grounded Emotional Support: An Interaction Design Framework and Its Multimodal Agentic Instantiation

用于文化根基型情感支持的诗意遗产:一种交互设计框架及其多模态智能体实例化

Yangming Zhang, Zhiqian Li, Bin Wu, Qi Li, Jie Xu, Yunpeng Song, Liang Zhao

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本研究提出一种将诗意传统转化为文化根基型情感支持交互媒介的设计框架,开发了基于LLM的多模态多智能体系统Poemithy,实验表明该系统可有效改善情绪等指标,多模态呈现能提升用户共鸣与参与度。

Comments 45 pages, 14 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07008 2026-08-25 cs.CV cs.LG 版本更新 70%

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints

不应学习的地方:基于子集归因约束的先验对齐训练以实现可靠的决策制定

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Kangwei Liu, Sanyi Zhang, Zhangcheng Wang, Shiming Liu, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) Communication University of China(中国传媒大学) Imperial College London(伦敦帝国学院) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络科学与技术学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出了一种基于归因的先验对齐方法,通过子集选择归因技术约束模型依赖于人类先验区域,从而提升决策的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23138 2026-08-25 cs.RO cs.AI cs.CV 新提交 62%

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Pointing-VLA:面向视觉-语言-动作操控的类型化空间定位接口

Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Pointing-VLA是基于Embodied-R1的类型化空间读出接口,可提升VLA模型的机器人操控性能,在多任务评估中实现SOTA表现,还能高效迁移至其他机器人系统并提升真实机器人自主成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20129 2026-08-25 cs.MA cs.CL cs.CV 版本更新 62%

Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

结合大语言模型常识推理能力的多智能体协同框架用于自动驾驶

Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 该研究针对自动驾驶中强化学习等方法的上下文推理缺陷,提出结合LLM常识推理的混合多智能体协同框架,经CARLA场景验证可保留结构化控制与安全机制,具备应用潜力。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23473 2026-08-25 cs.LG cs.AI 新提交 57%

MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters

MetaCaster:用于轻量级时间序列预测器小样本端到端学习的元调控优化智能体

ChengAo Shen, Wenchao Yu, Fangyu Wu, Dongjin Song, Hanghang Tong, Dongsheng Luo, Wei Cheng, Haifeng Chen, Jingchao Ni

机构 * University of Houston(休斯顿大学) NEC Labs(NEC实验室) University of Waterloo(滑铁卢大学) University of Connecticut(康涅狄格大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Singapore Management University(新加坡管理大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 针对资源受限场景下轻量级时间序列预测器小样本学习的困境,提出MetaCaster多智能体框架,可高效训练专用预测器,在18个数据集等实验中兼顾数据、计算效率与预测性能。

Comments Accepted by EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22449 2026-08-25 cs.RO cs.AI 新提交 57%

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

EMPIRE:将显式操作规划作为可学习中间表征用于自我中心视角下手运动预测

Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 该研究针对现有手运动预测方法忽略操作过程、梯度干扰的问题,提出两阶段框架EMPIRE,构建EMPIRE-651K数据集,实现了更优的手运动预测精度。

Comments 14 pages, 10 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21830 2026-08-25 cs.AI 新提交 57%

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

超越成功与失败:面向GUI智能体的长度感知对比学习

Chengyang Gu, Le Zhang, Jingbo Zhou, Yize Chen, Yu Shi, Siqi Bao, Zheng-Fan Wu, Hua Wu, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Baidu Inc.(百度公司) University of Alberta(阿尔伯塔大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 针对GUI智能体现有对比RLVR方法无法捕获轨迹细粒度质量差异的问题,提出LACL-GUI框架,引入轨迹级质量信号,在基准测试中实现性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08405 2026-08-25 cs.AI physics.flu-dyn 版本更新 57%

Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control

自进化科学智能体发现可泛化的物理推理流体控制

Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 提出一种由大语言模型驱动的自进化科学智能体工作流,通过迭代代码生成和物理仿真诊断,自动构建可解释的控制器,并在欠驱动双关节狗鲨游泳器目标到达任务中实现零样本泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01063 2026-08-25 cs.AI 版本更新 57%

MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

MindClaw: 用于精确干预的闭环具身心理状态推理

Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Wen-Huang Cheng, Jianlong Fu

机构 * Jilin University(吉林大学) Microsoft Asia(微软亚洲) National Taiwan University(国立台湾大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 提出MindClaw框架,通过闭环具身心理状态推理实现精确干预,结合多源输入、信念记忆、认知触发技能和动作生成,在动态环境中优化干预时机。

Comments Extended version of the CVPR 2026 paper *MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents*. This work is in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10169 2026-08-25 cs.AI cs.LG 版本更新 57%

MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction

MAVEN-T:用于实时多智能体轨迹预测的强化异构蒸馏

Wenchang Duan, Zhenguo Gao, Jinguo Xian, Yi Shi

机构 * School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院) Bio-X Institutes, Key Laboratory for the Genetics of Developmental and Neuropsychiatric Disorders, Shanghai Jiao Tong University(上海交通大学Bio-X研究院、发育与神经精神疾病遗传学重点实验室) Shanghai Key Laboratory of Psychotic Disorders, Brain Science and Technology Research Center, Shanghai Jiao Tong University(上海精神疾病重点实验室、脑科学与技术研究中心,上海交通大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 提出MAVEN-T框架,通过高容量教师模型和紧凑学生模型的异构蒸馏,结合强化学习优化,实现实时多智能体轨迹预测,在多个数据集上达到高精度与低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22702 2026-08-25 cs.HC 新提交 50%

AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations

AffAdapt:面向流畅对话的情感驱动型自适应AI角色

Nishanth Chidambaram, Kaustubh Paliwal, Kayla Hom, Shaoze Zhou, Chen Chen, Manas Satish Bedmutha, Nadir Weibel

专题命中 多模态Agent :multimodal(abstract)

AI总结 研究提出AffAdapt框架,整合多模块构建AI角色交互循环,实现流畅人机对话,在高风险对话场景验证其有效性,指出相关挑战并说明其应用场景。

Comments 3 pages, 2 figures, Adjunct Proceedings of the 39th Annual ACM Symposium on User Interface Software and Technology (UIST '26 Adjunct), Detroit, MI, USA

详情

展开后加载摘要…

URL PDF HTML 收藏