arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-11 至 2025-08-11 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2505.19015 2025-08-11 cs.CV cs.MM 84%

Can Multimodal Large Language Models Understand Spatial Relations?

Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou, Yinan Zou, Weiyan Zhang, Haiyun Jiang, Tong Ruan

机构 * School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China(信息科学与工程学院,东华大学,上海,中国) School of Computer Science, Fudan University, Shanghai, China(计算机科学学院,复旦大学,上海,中国)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures, published to ACL 2025

Journal ref In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 620-632, 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05684 2025-08-11 cs.CR cs.LG 82%

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

Junhao He, Tianyu Liu, Jingyuan Zhao, Benjamin Turner

机构 * Huaiyin Institute of Technology(淮阴职业技术学院) Universidad Autónoma de Asunción(阿斯unción自治大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09688 2025-08-11 cs.CV 79%

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

Jiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee, Hyunjung Shim

机构 * KAIST, Republic of Korea(韩国科学技术院) Samsung Electronics, Republic of Korea(三星电子)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06452 2025-08-11 cs.CV cs.LG 57%

TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation

Mattia Litrico, Mario Valerio Giuffrida, Sebastiano Battiato, Devis Tuia

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2508.06117 2025-08-11 cs.HC 78%

A Multimodal Framework for Understanding Collaborative Design Processes

Maurice Koch, Nelusa Pathmanathan, Daniel Weiskopf, Kuno Kurzhals

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted to IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21772 2025-08-11 cs.MM cs.AI 62%

Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline

Minwoo Oh, Minsu Park, Eunil Park

机构 * Department of MetaBioHealth, Sungkyunkwan University, Korea(韩国成均馆大学代谢生物健康系) Department of Applied Artificial Intelligence, Sungkyunkwan University, Korea(韩国成均馆大学应用人工智能系) Department of Computer Science and Engineering, Jaume I University, Spain(西班牙伊萨贝拉大学计算机科学与工程系)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted for publication at IJCAI 2025. 9 pages, 4 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06382 2025-08-11 cs.CV 57%

Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning

Xiangyu Wu, Feng Yu, Yang Yang, Jianfeng Lu

机构 * Nanjing University of Science

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05978 2025-08-11 cs.SD cs.AI cs.LG 57%

DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

Wei Chen, Binzhu Sha, Dan Luo, Jing Yang, Zhuo Wang, Fan Fan, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.AI

Comments Accepted by INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2508.05695 2025-08-11 cs.CR cs.LG 78%

MambaITD: An Efficient Cross-Modal Mamba Network for Insider Threat Detection

Kaichuan Kong, Dongjie Liu, Xiaobo Jin, Zhiying Li, Guanggang Geng, Jian Weng

机构 * College of Cyber Security(网络安全学院) Jinan University(济南大学) Department of Electrical and Electronic Engineering(电子与电气工程系) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 视频多模态 :cross-modal(title,abstract)

Comments Submitted to the 2025 IEEE International Conference on Data Mining (ICDM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15384 2025-08-11 cs.CV cs.AI 62%

MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies

Long Yang, Lianqing Zheng, Wenjin Ai, Minghao Liu, Sen Li, Qunshu Lin, Shengyu Yan, Jie Bai, Zhixiong Ma, Tao Huang, Xichan Zhu

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automobile, Chang'an University(长安大学汽车学院) School of Information and Electrical Engineering, Hangzhou City University(杭州城市学院信息与电气工程学院) College of Science and Engineering, James Cook University(詹姆斯库克大学科学与工程学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04592 2025-08-11 cs.LG cs.AI cs.CE cs.IR 57%

CAMEF: Causal-Augmented Multi-Modality Event-Driven Financial Forecasting by Integrating Time Series Patterns and Salient Macroeconomic Announcements

Yang Zhang, Wenbo Yang, Jun Wang, Qiang Ma, Jie Xiong

机构 * Southwestern University of Finance and Economics(西南财经大学) Kyoto Institute of Technology(京都技术大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted in SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13667 2025-08-11 cs.CV 57%

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,计算机学院,武汉大学) Hong Kong University of Science and Technology(香港科学与技术大学) Horizon Robotics

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10683 2025-08-11 cs.RO cs.AI 57%

Learning to Initialize Trajectory Optimization for Vision-Based Autonomous Flight in Unknown Environments

Yicheng Chen, Jinjie Li, Wenyuan Qin, Yongzhao Hua, Xiwang Dong, Qingdong Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to IROS 2025. Source code available

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2508.06154 2025-08-11 cs.IR cs.AI cs.MM 81%

Semantic Item Graph Enhancement for Multimodal Recommendation

Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Dusit Niyato, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06328 2025-08-11 cs.IR 78%

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation

Zhiyou Xiao, Qinhan Yu, Binghui Li, Geng Chen, Chong Chen, Wentao Zhang

专题命中 跨模态检索 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06104 2025-08-11 cs.CV 70%

MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment

Gui Zou, Chaofan Gan, Chern Hong Lim, Supavadee Aramvith, Weiyao Lin

机构 * Shanghai Jiao Tong University, China(上海交通大学) Monash University, Malaysia(墨尔本大学) Chulalongkorn University, Thailand(朱拉隆功大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICMEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14428 2025-08-11 cs.CV cs.LG q-bio.QM 57%

WildSAT: Learning Satellite Image Representations from Wildlife Observations

Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) GenBio AI University of Edinburgh(爱丁堡大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2508.05954 2025-08-11 cs.CV cs.AI cs.CL 85%

Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents

Han Lin, Jaemin Cho, Amir Zadeh, Chuan Li, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Lambda

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://bifrost-1.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06492 2025-08-11 cs.CV cs.CL 81%

Effective Training Data Synthesis for Improving MLLM Chart Understanding

Yuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li, Gaowen Liu, Ali Payani, Yuan-Sen Ting, Liang Zheng

机构 * Australian National University(澳大利亚国立大学) Ohio State University(俄亥俄州立大学) Cisco(思科公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025 (poster). 26 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03256 2025-08-11 cs.GR cs.CV 79%

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation

Xinyang Li, Gen Li, Zhihui Lin, Yichen Qian, GongXin Yao, Weinan Jia, Aowen Wang, Weihua Chen, Fan Wang

机构 * Xunguang Team, DAMO Academy, Alibaba Group(达摩院、阿里集团) Zhejiang University(浙江大学) Hupan Lab(虎盘实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06054 2025-08-11 eess.SP 78%

Multi-Modal Neural Radio Radiance Field for Localized Statistical Channel Modelling

Yiheng Wang, Shutao Zhang, Ye Xue, Tsung-Hui Chang

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08501 2025-08-11 cs.LG 50%

Learning to Match Unpaired Data with Minimum Entropy Coupling

Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

机构 * Department of Data Science, EURECOM, France(数据科学系,EURECOM,法国)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 6 篇

2508.06009 2025-08-11 cs.CV 79%

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models

Jun Feng, Zixin Wang, Zhentao Zhang, Yue Guo, Zhihan Zhou, Xiuyi Chen, Zhenyang Li, Dawei Yin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05669 2025-08-11 cs.IR cs.AI cs.CL cs.CV cs.LG 67%

Fine-Tuning Vision-Language Models for Markdown Conversion of Financial Tables in Malaysian Audited Financial Reports

Jin Khye Tan, En Jun Choong, Ethan Jeremiah Chitty, Yan Pheng Choo, John Hsin Yang Wong, Chern Eu Cheah

专题命中 多模态评测 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 14 figures, 5 tables. Evaluation code (LLM-as-a-judge and Markdown TEDS) is available at https://github.com/jinkhye/MyFinMarkdown. The development dataset and evaluation benchmark are available on Hugging Face at https://huggingface.co/datasets/jinkhye/MyFinMarkdown-sample and https://huggingface.co/datasets/jinkhye/MyFinMarkdown-bench respectively

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09105 2025-08-11 cs.CV cs.AI cs.CL cs.LG 67%

INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance

Chenwei Lin, Hanjia Lyu, Xian Xu, Jiebo Luo

机构 * School of Computer Science Fudan University(复旦大学计算机科学学院) Department of Computer Science University of Rochester(罗切斯特大学计算机科学系) School of Economics Fudan University(复旦大学经济学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear in the International Conference on Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06345 2025-08-11 cs.CL cs.AI cs.GR cs.LG 62%

Harnessing Adaptive Topology Representations for Zero-Shot Graph Question Answering

Yanbin Wei, Jiangyue Yan, Chun Kang, Yang Chen, Hua Liu, James T. Kwok, Yu Zhang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06072 2025-08-11 cs.CV cs.AI 62%

Can Large Models Fool the Eye? A New Turing Test for Biological Animation

Zijian Chen, Lirong Deng, Zhengyu Chen, Kaiwei Zhang, Qi Jia, Yuan Tian, Yucheng Zhu, Guangtao Zhai

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09505 2025-08-11 cs.RO 50%

TruckV2X: A Truck-Centered Perception Dataset

Tenghui Xie, Zhiying Song, Fuxi Wen, Jun Li, Guangzhao Liu, Zijian Zhao

专题命中 多模态评测 :multi-modal(abstract)

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 9, pp. 9312-9319, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 5 篇

2405.00522 2025-08-11 econ.GN cs.CE cs.CL cs.CR q-fin.CP q-fin.EC 79%

DAM: A Universal Dual Attention Mechanism for Multimodal Timeseries Cryptocurrency Trend Forecasting

Yihang Fu, Mingyu Zhou, Luyao Zhang

机构 * Duke Kunshan University(杜克昆山大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Journal ref Proc. IEEE Int. Conf. Metaverse Computing Networking and Applications (MetaCom), pp. 73-80, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06342 2025-08-11 cs.CV cs.SI 57%

Street View Sociability: Interpretable Analysis of Urban Social Behavior Across 15 Cities

Kieran Elrod, Katherine Flanigan, Mario Bergés

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏