arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-10 至 2025-12-10 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 15 篇

2505.07007 2025-12-10 cs.CV 57%

MELLM: A Flow-Guided Large Language Model for Micro-Expression Understanding

MELLM:一种面向微表情理解的流引导大型语言模型

Sirui Zhao, Zhengye Zhang, Shifeng Liu, Xinglong Mao, Shukang Yin, Chaoyou Fu, Tong Xu, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) Nanjing University(南京大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 MELLM通过结合光学流敏感性与LLM推理能力,首次实现对微表情的全面理解,显著提升微表情识别的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12695 2025-12-10 cs.RO cs.SY eess.SY 50%

CDKFormer: Contextual Deviation Knowledge-Based Transformer for Long-Tail Trajectory Prediction

CDKFormer: 基于上下文偏差知识的长尾轨迹预测变换器

Yuansheng Lian, Ke Zhang, Meng Li

专题命中 多模态评测 :multimodal(abstract)

AI总结 CDKFormer通过引入上下文偏差知识和双查询解码器,提升自动驾驶车辆在长尾轨迹预测中的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07912 2025-12-10 q-bio.QM 50%

A Semi-Supervised Inf-Net Framework for CT-Based Lung Nodule Analysis with a Conceptual Extension Toward Genomic Integration

一种基于CT的肺结节分析的半监督Inf-Net框架及其向基因组整合的概念扩展

Fateme Mobini, Mohammad Reza Hedyehzadeh, Mahdi Yousefi

专题命中 多模态评测 :multimodal(abstract)

AI总结 本研究提出一种半监督Inf-Net框架,通过整合基因组数据提升CT肺结节分析精度,实验表明其在多模态诊断中具有潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 6 篇

2512.08629 2025-12-10 cs.AI cs.CV cs.HC 84%

See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm

See-Control: 一种多模态代理框架用于智能手机与机械臂的交互

Haoyu Zhao, Weizhong Ding, Yuhao Yang, Zheng Tian, Linyi Yang, Kun Shao, Jun Wang

机构 * University College London(伦敦大学学院) Imperial College London(伦敦帝国学院) Huawei Noah’s Ark Lab(华为诺亚实验室) ShanghaiTech University(上海科技大学) Southern University of Science(南方科技大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 See-Control通过直接与机械臂交互实现智能手机操作,无需ADB或系统后台访问,提供了一个平台无关的多模态代理框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02530 2025-12-10 cs.AI 79%

Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration

Aetheria:基于多智能体辩论与协作的多模态可解释内容安全框架

Yuxiang He, Jian Zhao, Yuchen Yuan, Tianle Zhang, Wei Cai, Haojie Cheng, Ziyan Shi, Ming Zhu, Haichuan Tang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Sichuan University(四川大学) Peking University(北京大学) Beijing Jiaotong University(北京交通大学) Harbin Institute of Technology(哈尔滨工业大学) China Railway Rolling Stock Corporation Limited(中国铁路滚动股票有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 Aetheria通过多智能体辩论与协作机制,实现多模态内容安全的可解释性与高准确性,提升可信AI审查水平。

Comments https://github.com/Herrieson/Aetheria

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08769 2025-12-10 cs.AI 57%

A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows

生产级智能体AI工作流设计、开发与部署的实用指南

Eranga Bandara, Ross Gore, Peter Foytik, Sachin Shetty, Ravi Mukkamala, Abdul Rahman, Xueping Liang, Safdar H. Bouk, Amin Hass, Sachini Rajapakse, Ng Wee Keong, Kasun De Zoysa, Aruna Withanage, Nilaan Loganathan

机构 * Old Dominion University(旧 Dominion 大学) Florida International University(佛罗里达国际大学) Nanyang Technological University(南洋理工大学) University of Colombo(科伦坡大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一套生产级智能体AI工作流的设计、开发和部署的最佳实践,涵盖架构设计、多智能体模式、模型上下文协议及环境感知部署策略,旨在提升系统的可靠性与可维护性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08476 2025-12-10 cs.RO 50%

A Multi-Agent LLM Framework for Design Space Exploration in Autonomous Driving Systems

多智能体大语言模型框架用于自动驾驶系统的设计空间探索

Po-An Shih, Shao-Hua Wang, Yung-Che Li, Chia-Heng Tu, Chih-Han Chang

机构 * National Cheng Kung University(国立成功大学) Safeware Technology Inc.(Safeware技术公司)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出基于多智能体大语言模型的自动驾驶系统设计空间探索框架,通过多模态推理和3D模拟工具自动化解析执行输出,提升设计效率和成本效益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08186 2025-12-10 cs.RO 50%

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation

放慢脚步,快速移动:一种通用视觉-语言导航的双系统基础模型

Meng Wei, Chenyang Wan, Jiaqi Peng, Xiqian Yu, Yuqiang Yang, Delin Feng, Wenzhe Cai, Chenming Zhu, Tai Wang, Jiangmiao Pang, Xihui Liu

机构 * Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 DualVLN通过双系统架构,结合高层推理与低层动作执行,提升视觉-语言导航的泛化能力和动态环境适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23823 2025-12-10 cs.RO 50%

Control Your Robot: A Unified System for Robot Control and Policy Deployment

掌控你的机器人:一种统一的机器人控制与策略部署系统

Tian Nian, Weijie Ke, Shaolong Zhu, Bingshan Hu

机构 * ScaleLab, Shanghai Jiao Tong University(上海交通大学ScaleLab) University of Shanghai for Science and Technology(上海科学技术大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Control Your Robot提出了一种统一的机器人控制与策略部署框架,通过模块化设计和标准化流程,实现跨平台的数据收集与策略学习,提升机器人学习的可扩展性和可重复性。

Comments Code: https://github.com/Tian-Nian/control_your_robot

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 6 篇

2508.07871 2025-12-10 cs.CV 85%

CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning

CATP:面向高效增强多模态上下文学习的上下文自适应令牌剪枝

Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 CATP通过上下文自适应令牌剪枝方法,提升多模态上下文学习的效率和性能,减少冗余令牌带来的影响。

Comments 14 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08820 2025-12-10 cs.CV cs.AI 81%

Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning

无需训练的双双曲适配器用于更高效的跨模态推理

Yi Zhang, Chun-Wun Cheng, Junyi He, Ke Yu, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Aviles-Rivero

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Department of Electrical and Electronic Engineering, Southern University of Science and Technology(南方科技大学电子与电气工程系) Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学应用数学与理论物理系) Yau Mathematical Sciences Center, Tsinghua University(清华大学应用数学中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出无需训练的双双曲适配器方法,通过双曲空间嵌入提升跨模态推理性能,实现更高效的领域泛化和少样本识别。

Comments Accepted in IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08135 2025-12-10 cs.CV 79%

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning

CVP:基于中央-外围视觉的多模态模型用于空间推理

Zeyuan Chen, Xiang Zhang, Haiyang Xu, Jianwen Xie, Zhuowen Tu

机构 * UC San Diego(圣迭戈大学) Lambda, Inc(Lambda公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 CVP通过结合中央视觉和外围视觉的启发,提出了一种多模态模型,以提升复杂3D环境的空间推理能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08071 2025-12-10 cs.LG 78%

CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification

CAMO:基于因果引导的对抗多模态领域泛化用于危机分类

Pingchuan Ma, Chengshuai Zhao, Bohan Jiang, Saketh Vishnubhatla, Ujun Jeong, Alimohammad Beigi, Adrienne Raglin, Huan Liu

机构 * Arizona State University(亚利桑那州立大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 CAMO通过结合对抗去耦与统一表示学习,解决多模态危机分类中的领域泛化问题,提升模型在未见灾害场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08374 2025-12-10 cs.CV 70%

The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss

看不见的偏见:预规范MLLM中规范差异如何导致视觉信息丢失

Bozhou Li, Xinda Xue, Sihan Yang, Yang Shi, Xinlong Chen, Yushuo Guan, Yuanxing Zhang, Wentao Zhang

机构 * Peking University(北京大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Xi’an Jiaotong University(西安交通大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文揭示了预规范MLLM中规范差异导致的视觉信息丢失问题,并提出通过插入层规范层来解决这一问题,提升模型整体能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08572 2025-12-10 cs.CV 57%

From Cells to Survival: Hierarchical Analysis of Cell Inter-Relations in Multiplex Microscopy for Lung Cancer Prognosis

从细胞到生存:多乘 microscopy 中细胞互相关系的分层分析用于肺癌预后

Olle Edgren Schüllerqvist, Jens Baumann, Joakim Lindblad, Love Nordling, Artur Mezheyeuski, Patrick Micke, Nataša Sladoje

机构 * Department of Information Technology, SciLifeLab, Uppsala University(信息科技系、SciLifeLab、乌普萨拉大学) PAICON GmbH(PAICON公司) Department of Immunology, Genetics and Pathology, Uppsala University(免疫学、遗传学和病理学系、乌普萨拉大学) Molecular Oncology Group, Vall d'Hebron Institute of Oncology(分子肿瘤学组、瓦伦·德·霍布伦肿瘤研究所) Vall d'Hebron Institute of Research(瓦伦·德·霍布伦研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 HiGINE通过多光谱免疫荧光图像分析肿瘤微环境,利用分层图方法预测肺癌患者生存并提升风险分层。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 5 篇

2512.08889 2025-12-10 cs.CV cs.AI 76%

No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers

无标签,无问题:利用多模态验证器训练视觉推理器

Damiano Marsili, Georgia Gkioxari

机构 * California Institute of Technology(加州理工学院)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出无需标注的视觉推理训练框架,结合AI驱动的验证器提升推理与定位能力,超越现有开源和专有模型。

Comments Project webpage: https://glab-caltech.github.io/valor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08567 2025-12-10 cs.LG cs.AI 57%

A Hybrid Model for Stock Market Forecasting: Integrating News Sentiment and Time Series Data with Graph Neural Networks

一种股票市场预测的混合模型:整合新闻情感与时间序列数据与图神经网络

Nader Sadek, Mirette Moawad, Christina Naguib, Mariam Elzahaby

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种结合新闻情感与时间序列数据的混合模型,利用图神经网络提升股票市场预测性能,实验显示GNN在准确率和精度上均优于LSTM基线。

Comments 11 pages, 6 figures. Published in the Proceedings of the 5th International Conference on Artificial Intelligence Research (ICAIR 2025). Published version available at: https://papers.academic-conferences.org/index.php/icair/article/view/4294

Journal ref Proceedings of the 5th International Conference on AI Research (ICAIR 2025), Vol. 5, No. 1, pp. 452-462 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08262 2025-12-10 cs.CV cs.RO 57%

RLCNet: An end-to-end deep learning framework for simultaneous online calibration of LiDAR, RADAR, and Camera

RLCNet: 一种用于同时在线校准激光雷达、雷达和摄像头的端到端深度学习框架

Hafeez Husain Cholakkal, Stefano Arrigoni, Francesco Braghin

机构 * Department of Mechanical Engineering, Politecnico di Milano(机械工程系,米兰理工学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 RLCNet通过端到端深度学习框架实现激光雷达、雷达和摄像头的同时在线校准,提升自动驾驶系统在动态环境中的感知可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08221 2025-12-10 cs.CV 57%

VisKnow: Constructing Visual Knowledge Base for Object Understanding

VisKnow: 构建视觉知识库以实现物体理解

Ziwei Yao, Qiyang Wan, Ruiping Wang, Xilin Chen

机构 * Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室、计算技术研究所)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 VisKnow通过构建视觉知识库,利用多模态数据提升物体理解能力,包含AnimalKB案例研究,用于零样本识别和知识图谱完成等任务。

Comments 16 pages, 12 figures, 7 tables. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08316 2025-12-10 astro-ph.IM astro-ph.CO 50%

PolySwyft: sequential simulation-based nested sampling

PolySwyft:基于序列模拟的嵌套采样

Kilian H. Scheutwinkel, Will Handley, Christoph Weniger, Eloy de Lera Acedo

专题命中 其他多模态 :multimodal(abstract)

AI总结 PolySwyft结合嵌套采样和神经比率估计,通过序列模拟方法高效推断复杂后验分布,适用于似然不可行但正向模拟器可用的场景。

详情

展开后加载摘要…

URL PDF HTML 收藏