arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-22 至 2025-08-22 共收录 42 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 5 篇

2508.15164 2025-08-22 cs.CL 57%

ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following

Seungmin Han, Haeun Kwon, Ji-jun Park, Taeyang Yoon

机构 * Dongguk University(东国大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21969 2025-08-22 cs.RO cs.AI 57%

Embodied Long Horizon Manipulation with Closed-loop Code Generation and Incremental Few-shot Adaptation

Yuan Meng, Xiangtong Yao, Haihui Ye, Yirui Zhou, Shengqiang Zhang, Zhenguo Sun, Xukun Li, Zhenshan Bing, Alois Knoll

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments update ICRA 6 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15493 2025-08-22 cond-mat.mtrl-sci 50%

Attention-Based Explainability for Structure-Property Relationships

Boris N. Slautin, Utkarsh Pratiush, Yongtao Liu, Hiroshi Funakubo, Vladimir V. Shvartsman, Doru C. Lupascu, Sergei V. Kalinin

专题命中 多模态Agent :multimodal(abstract)

Comments 34 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 4 篇

2508.15505 2025-08-22 cs.CV 79%

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang, Yuwei Wang, Yanyan Wei, Fei Wang

机构 * Key Laboratory of Opto-Electronic Information Science and Technology of Jiangxi Province, Nanchang Hangkong University(江西省光电信息科学与技术重点实验室,南昌航空大学) ReLER, CCAI, Zhejiang University(ReLER、CCAI、浙江大学) College of Engineering, Anhui Agricultural University(安徽农业大学工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13792 2025-08-22 cs.LG cs.AI cs.RO 79%

Continual Learning for Multimodal Data Fusion of a Soft Gripper

Nilay Kushawaha, Egidio Falotico

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted in Wiley Advanced Robotics Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14485 2025-08-22 cs.IR 78%

Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion

Moyu Zhang, Yongxiang Tang, Yujun Jin, Jinxin Hu, Yu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025, 11 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09588 2025-08-22 cs.CV cs.AI 62%

TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting

Zhicong Wu, Hongbin Xu, Gang Xu, Ping Nie, Zhixin Yan, Jinkai Zheng, Liangqiong Qu, Ming Li, Liqiang Nie

机构 * Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究所) Guangdong Laboratory of Artificial Intelligence(广东省人工智能与数字经济实验室) Peking University(北京大学) School of Future Technology, South China University of Technology(华南理工大学未来技术学院) Hangzhou Dianzi University(杭州电子科技大学) Central Laboratory of Lishui Hospital of Wenzhou Medical University, The First Affiliated Hospital of Lishui University, Lishui People's Hospital(丽水市人民医院中央实验室、丽水大学第一附属医院、丽水人民医院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 5 篇

2506.05182 2025-08-22 cs.IR 67%

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza, Yifan Nie, Jimmy Lin

专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract)

Comments 15 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15752 2025-08-22 cs.HC cs.AI cs.CV 62%

"Does the cafe entrance look accessible? Where is the door?" Towards Geospatial AI Agents for Visual Inquiries

Jon E. Froehlich, Jared Hwang, Zeyu Wang, John S. O'Meara, Xia Su, William Huang, Yang Zhang, Alex Fiannaca, Philip Nelson, Shaun Kane

机构 * University of Washington(华盛顿大学) Google Research(谷歌研究) UCLA(加州大学洛杉矶分校) Google DeepMind(谷歌DeepMind)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to the ICCV'25 Workshop "Vision Foundation Models and Generative AI for Accessibility: Challenges and Opportunities"

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06911 2025-08-22 cs.LG cs.AI 57%

MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning

Lishan Yang, Wei Emma Zhang, Quan Z. Sheng, Lina Yao, Weitong Chen, Ali Shakeri

机构 * The University of Adelaide(阿德莱德大学) Macquarie University(麦考瑞大学) CSIRO’s Data 61 and The University of New South Wales(CSIRO的数据61与新南威尔士大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04742 2025-08-22 cs.CY cs.AI 57%

A Case for Specialisation in Non-Human Entities

El-Mahdi El-Mhamdi, Lê-Nguyên Hoang, Mariame Tighanimine

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15638 2025-08-22 quant-ph 50%

Quantum Co-Magnetometer Using Diamond Nitrogen-Vacancy Centers and Rubidium Cells

Ittai Shalev, Kfir Levi, Rotem Malkinson, Amir Hen, Liron Stern, Nir Bar-Gill

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏