arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-17 至 2025-09-17 共收录 53 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2507.08679 2025-09-17 cs.CV 79%

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee, Kousik Dasgupta, Subarna Tripathi

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) Intel Labs(英特尔实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13282 2025-09-17 cs.CL cs.CV cs.LG 62%

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain, Leonid Sigal, Giuseppe Carenini

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10105 2025-09-17 cs.CV cs.CL 62%

VARCO-VISION-2.0 Technical Report

Young-rok Cha, Jeongho Ju, SunYoung Park, Jong-Hyeon Lee, Younghyun Yu, Youngjune Kim

机构 * NC AI

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13021 2025-09-17 cs.CL cs.CV 62%

Dynamic Relation Inference via Verb Embeddings

Omri Suissa, Muhiim Ali, Ariana Azarbal, Hui Shen, Shekhar Pradhan

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15244 2025-09-17 cs.CV cs.AI 62%

Adversarial Prompt Distillation for Vision-Language Models

Lin Luo, Xin Wang, Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12492 2025-09-17 cs.CV 57%

Evaluating Robustness of Vision-Language Models Under Noisy Conditions

Purushoth, Alireza

机构 * University of Nevada Reno(内华达大学拉斯维加斯分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 1 篇

2507.02844 2025-09-17 cs.CV cs.CL cs.CR 62%

Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection

Ziqi Miao, Yi Ding, Lijun Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to EMNLP 2025 (Main). 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2509.12893 2025-09-17 cs.CV 83%

MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization

Yiyi Zhang, Yuchen Yuan, Ying Zheng, Jialun Pei, Jinpeng Li, Zheng Li, Pheng-Ann Heng

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12269 2025-09-17 cs.LG cs.IR 71%

Research on Short-Video Platform User Decision-Making via Multimodal Temporal Modeling and Reinforcement Learning

Jinmeiyang Wang, Jing Dong, Li Zhou

专题命中 视频多模态 :multimodal(title)

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12231 2025-09-17 cs.DC 67%

Research on fault diagnosis and root cause analysis based on full stack observability

Jian Hou

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12876 2025-09-17 cs.CL cs.MM 62%

Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents

Fuyu Xing, Zimu Wang, Wei Wang, Haiyang Zhang

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL、cs.MM

Comments Accepted at INLG 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12741 2025-09-17 cs.RO cs.AI cs.LG 57%

Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions

Alexis Yihong Hao, Yufei Wang, Navin Sriram Ravie, Bharath Hegde, David Held, Zackory Erickson

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Department of Engineering Design, Indian Institute of Technology, Madras(印度理工学院Madras分校工程设计系)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2509.12600 2025-09-17 cs.LG cs.AI q-bio.QM 88%

A Multimodal Foundation Model to Enhance Generalizability and Data Efficiency for Pan-cancer Prognosis Prediction

Huajun Zhou, Fengtao Zhou, Jiabo Ma, Yingxue Xu, Xi Wang, Xiuming Zhang, Li Liang, Zhenhui Li, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Hong Kong University of Science and Technology(香港科学与技术大学) Department of Pathology(病理学系) School of Medicine(医学院) Zhejiang University(浙江大学) Nanfang Hospital and School of Basic Medical Sciences(南方医科大学基础医学系) Southern Medical University(南方医学院) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理重点实验室) Jinfeng Laboratory(金凤实验室) Department of Radiology(放射科) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders(神经系统疾病国家重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12653 2025-09-17 cs.CV cs.AI 84%

Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations

Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong

机构 * Hefei University of Technology(合肥工业大学) University of Trento(特伦托大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01275 2025-09-17 cs.AI 83%

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou, silvio savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles

机构 * Salesforce AI Reseach(Salesforce人工智能研究院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12994 2025-09-17 cs.CL 70%

SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data

Jian Gao, Fufangchen Zhao, Yiyang Zhang, Danfeng Yan

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13175 2025-09-17 cs.CV 57%

More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era

Yingtai Li, Haoran Lai, Xiaoqian Zhou, Shuai Ming, Wenxin Ma, Wei Wei, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei Anhui, 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advance Research, USTC, Suzhou Jiangsu, 215123, China The First Affiliated Hospital of USTC, Division of Life Sciences Medicine, USTC, Hefei Anhui, 230001, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21875 2025-09-17 cs.AI 57%

Tiny-BioMoE: a Lightweight Embedding Model for Biosignal Analysis

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2509.12883 2025-09-17 cs.CV 83%

Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder

Qifei Jia, Yu Liu, Yajie Chai, Xintong Yao, Qiming Lu, Yasen Zhang, Runyu Shi, Ying Huang, Guoquan Zhang

机构 * Xiaomi Corporation Beijing, China(小米公司北京)

专题命中 多模态生成 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09315 2025-09-17 cs.RO cs.CV cs.LG 79%

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving

Xuefeng Jiang, Yuan Ma, Pengxiang Li, Leimeng Xu, Xin Wen, Kun Zhan, Zhongpu Xia, Peng Jia, Xianpeng Lang, Sheng Sun

机构 * LiAuto Inc(LiAuto公司) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动研究所)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12888 2025-09-17 cs.CV cs.AI 62%

Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing

Weiming Chen, Zhihan Zhu, Yijia Wang, Zhihai He

机构 * Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12534 2025-09-17 eess.IV cs.AI cs.CV 62%

DeepEyeNet: Generating Medical Report for Retinal Images

Jia-Hong Huang

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments The paper is accepted by the Conference on Information and Knowledge Management (CIKM), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13235 2025-09-17 cs.AI 57%

A Scenario-Driven Cognitive Approach to Next-Generation AI Memory

Linyue Cai, Yuyang Cheng, Xiaoding Shao, Huiming Wang, Yong Zhao, Wei Zhang, Kang Li

机构 * School of Computer Science, Sichuan University(四川大学计算机学院) School of Cyber Science and Engineering, Sichuan University(四川大学网络科学与工程学院) West China Biomedical Big Data Center, Sichuan University West China Hospital(四川大学华西生物医学大数据中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2509.12060 2025-09-17 cs.AI 89%

When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models

Wei Cai, Shujuan Liu, Jian Zhao, Ziyan Shi, Yusheng Zhao, Yuchen Yuan, Tianle Zhang, Chi Zhang, Xuelong Li

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13234 2025-09-17 cs.AI cs.CV cs.HC 84%

Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

Nadim Barakat, William Lotter

机构 * Dana-Farber Cancer Institute & Tufts University School of Medicine(达纳-法伯癌症研究所及塔夫茨大学医学院) Dana-Farber Cancer Institute Brigham and Women’s Hospital & Harvard Medical School(达纳-法伯癌症研究所布里特妇女医院及哈佛医学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13289 2025-09-17 cs.CV eess.IV 79%

Image Realness Assessment and Localization with Multimodal Features

Lovish Kaushik, Agnij Biswas, Somdyuti Paul

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈拉格普)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12963 2025-09-17 cs.CV cs.LG 79%

MMMS: Multi-Modal Multi-Surface Interactive Segmentation

Robin Schön, Julian Lorenz, Katja Ludwig, Daniel Kienzle, Rainer Lienhart

机构 * Fakultät für Angewandte Informatik University of Augsburg(应用信息学院乌尔姆大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 19 pages, 11 figures, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12750 2025-09-17 cs.CV 79%

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

Rishab Parthasarathy, Jasmine Collins, Cory Stephenson

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 9 figures, 3 tables; appendix 16 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12287 2025-09-17 eess.IV cs.CV cs.LG 79%

Enhancing Radiographic Disease Detection with MetaCheX, a Context-Aware Multimodal Model

Nathan He, Cody Chen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments All authors contributed equally, 5 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19075 2025-09-17 cs.CV 79%

HoloDx: Knowledge- and Data-Driven Multimodal Diagnosis of Alzheimer's Disease

Qiuhui Chen, Jintao Wang, Gang Wang, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Department of Neurology, Renji Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属仁济医院神经内科)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Medical Imaging (TMI)

详情

展开后加载摘要…

URL PDF HTML 收藏