arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-03 至 2025-11-03 共收录 47 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2510.27443 2025-11-03 cs.LG 78%

MVeLMA: Multimodal Vegetation Loss Modeling Architecture for Predicting Post-fire Vegetation Loss

Meenu Ravi, Shailik Sarkar, Yanshen Sun, Vaishnavi Singh, Chang-Tien Lu

机构 * Georgetown University(乔治城大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted for 2025 ACM SIGSPATIAL conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27363 2025-11-03 cs.AI 70%

ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use

Mengjie Deng, Guanting Dong, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17336 2025-11-03 cs.MM cs.CL cs.CV 67%

Mano Technical Report

Tianyu Fu, Anyang Su, Chenxu Zhao, Hanning Wang, Minghui Wu, Zhe Yu, Fei Hu, Mingjia Shi, Wei Dong, Jiayao Wang, Yuyang Chen, Ruiyang Yu, Siran Peng, Menglin Li, Nan Huang, Haitian Wei, Jiawei Yu, Yi Xin, Xilin Zhao, Kai Gu, Ping Jiang, Sifan Zhou, Shuo Wang

机构 * Mininglamp Technology(Mininglamp技术)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27210 2025-11-03 cs.AI cs.CV 62%

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

Tao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu, Xuming He, Bai Song

机构 * ShanghaiTech University(上海科技大学) ByteDance(字节跳动) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程研究中心)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments Published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02652 2025-11-03 cs.AI cs.CL cs.IR 62%

HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search

Jiajie Jin, Xiaoxi Li, Guanting Dong, Yuyao Zhang, Yutao Zhu, Yang Zhao, Hongjin Qian, Zhicheng Dou

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13178 2025-11-03 cs.CR cs.AI cs.RO 57%

SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents

Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, Siheng Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Shanghai Jiao Tong University(上海交通大学) Department of Computer Science(计算机科学系) University of Georgia(佐治亚大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态Agent :image-text(abstract);分类 cs.AI

Comments 28 pages, 19 tables, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09255 2025-11-03 cs.HC 50%

Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents

Geonsun Lee, Min Xia, Nels Numan, Xun Qian, David Li, Yanhe Chen, Achin Kulshrestha, Ishan Chatterjee, Yinda Zhang, Dinesh Manocha, David Kim, Ruofei Du

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 7 篇

2510.27091 2025-11-03 cs.LG cs.AI quant-ph 83%

QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis

Yiwei Chen, Kehuan Yan, Yu Pan, Daoyi Dong

机构 * School of Engineering, Yunnan University(云南大学工程学院) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Institute of Cyber-Systems and Control, College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院智能系统与控制研究所) Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27208 2025-11-03 cs.CV cs.AI 81%

Multi-Modal Feature Fusion for Spatial Morphology Analysis of Traditional Villages via Hierarchical Graph Neural Networks

Jiaxin Zhang, Zehong Zhu, Junye Deng, Yunqin Li, and Bowen Wang

机构 * Architecture and Design College, Nanchang University(南昌大学建筑与设计学院) SANKEN, The University of Osaka(大阪大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27166 2025-11-03 cs.CV 79%

M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar

Xiaozhi Li, Huijun Di, Jian Li, Feng Liu, Wei Liang

机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究所,信息电子学院,北京理工大学) Key Laboratory of Electronic and Information Technology in Satellite Navigation, Ministry of Education(卫星导航电子信息技术重点实验室,教育部) School of Computer Science, Beijing Institute of Technology(计算机学院,北京理工大学) Beijing Racobit Electronic Information Technology Co., Ltd.(北京Racobit电子信息技术有限公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16026 2025-11-03 cs.LG stat.AP 79%

A tutorial on discovering and quantifying the effect of latent causal sources of multimodal EHR data

Marco Barbero-Mota, Eric V. Strobl, John M. Still, William W. Stead, Thomas A. Lasko

机构 * Department of Biomedical Informatics Vanderbilt University Medical Center(生物医学信息学系范德堡大学医学中心) Department of Biomedical Informatics University of Pittsburgh(生物医学信息学系匹兹堡大学) Departments of Medicine & Biomedical Informatics Vanderbilt University Medical Center(医学与生物医学信息学系范德堡大学医学中心) Departments of Biomedical Informatics & Computer Science Vanderbilt University Medical Center & Vanderbilt University(生物医学信息学与计算机科学系范德堡大学医学中心及范德堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at the 1st Multimodal Representation Learning for Healthcare EurIPS 2025 Workshop (https://multimodal-rep-learning-for-health.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27674 2025-11-03 physics.ed-ph 50%

Teaching competencies in physics for engineering education: A qualitative analysis from teaching practice

Vanessa Cruz Molina, Daniel Sanchez Guzman, Teodoro Rivera Montalvo, Ricardo Garcia-Salcedo

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26886 2025-11-03 cond-mat.mtrl-sci 50%

MaterialsGalaxy: A Platform Fusing Experimental and Theoretical Data in Condensed Matter Physics

Tiannian Zhu, Zhong Fang, Quansheng Wu, Hongming Weng

专题命中 多模态训练与对齐 :cross-modal(abstract)

Comments 18 pages, 5 figures. Accepted for publication in Chinese Physics B (24 October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10650 2025-11-03 stat.ML cs.LG stat.AP stat.CO stat.ME 50%

Generative Adversarial Networks for High-Dimensional Item Factor Analysis: A Deep Adversarial Learning Algorithm

Nanyu Luo, Feng Ji

机构 * University of Toronto(多伦多大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 3 篇

2306.03833 2025-11-03 cs.LG 78%

Decoding Virtual Healthcare Success through Knowledge-Aware and Multimodal Predictive Modeling

Shuang Geng, Wenli Zhang, Jiaheng Xie, Gemin Liang, Ben Niu, Sudha Ram

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27257 2025-11-03 cs.DC 50%

Synergistic Tensor and Pipeline Parallelism

Mengshi Qi, Jiaxuan Peng, Jie Zhang, Juan Zhu, Yong Li, Huadong Ma

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27040 2025-11-03 eess.SP cs.LG 50%

GeoPep: A geometry-aware masked language model for protein-peptide binding site prediction

Dian Chen, Yunkai Chen, Tong Lin, Sijie Chen, Xiaolin Cheng

专题命中 其他多模态 :multimodal(abstract)

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏