arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-19 至 2025-08-19 共收录 62 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 9 篇

2507.09693 2025-08-19 cs.CV 57%

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

Jiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang, Jianpeng Chen, Xusen Hei, Jiayuan Xie, Yi Cai, Qing Li

机构 * South China University of Technology(南方科技大学) The Hong Kong Polytechnic University(香港理工大学) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12701 2025-08-19 eess.SY cs.SY 50%

Deadline-Aware Bandwidth Allocation for Semantic Generative Communication with Diffusion Models

Jinhyuk Choi, Jihong Park, Seungeun Oh, Seong-Lyun Kim

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12629 2025-08-19 cs.LG q-bio.BM 50%

FlowMol3: Flow Matching for 3D De Novo Small-Molecule Generation

Ian Dunn, David R. Koes

机构 * Department of Computational and Systems Biology, University of Pittsburgh(计算与系统生物学系,匹兹堡大学)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 11 篇

2508.12591 2025-08-19 cs.CL cs.AI cs.SD 86%

Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning

Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen

机构 * National Taiwan Normal University(台湾国立正常大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted at IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12291 2025-08-19 cs.AI 83%

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

Xuming He, Zhiyuan You, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Wenlong Zhang, Lei Bai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ZheJiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Center for Earth System Modeling and Prediction of China Meteorological Administration(中国气象局地球系统模拟与预测中心)

专题命中 多模态评测 :multi-modal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11654 2025-08-19 eess.SP cs.CV 79%

Data-driven RF Tomography via Cross-modal Sensing and Continual Learning

Yang Zhao, Tao Wang, Said Elhadi

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 6 pages, 4 figures, to be published in IEEE AVSS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11738 2025-08-19 cs.CY cs.AI cs.CV 73%

Artificial Intelligence in Rural Healthcare Delivery: Bridging Gaps and Enhancing Equity through Innovation

Kiruthika Balakrishnan, Durgadevi Velusamy, Hana E. Hinkle, Zhi Li, Karthikeyan Ramasamy, Hikmat Khan, Srini Ramaswamy, Pir Masoom Shah

机构 * University of Illinois College of Medicine Rockford(伊利诺伊大学罗克福德医学院) National Center for Rural Health Professions(农村健康专业中心) Sri Sivasubramaniya Nadar College of Engineering(塞维萨布拉曼尼亚那德尔工程学院) M. Kumarasamy College of Engineering(M. Kumarasamy 工程学院) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医疗中心) iWorks Corporation(iWorks 公司) Central South University(中南大学) Bacha Khan University(巴哈汗大学)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11728 2025-08-19 cs.CV cs.AI 62%

UniDCF: A Foundation Model for Comprehensive Dentocraniofacial Hard Tissue Reconstruction

Chunxia Ren, Ning Zhu, Yue Lai, Gui Chen, Ruijie Wang, Yangyi Hu, Suyao Liu, Shuwen Mao, Hong Su, Yu Zhang, Li Xiao

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12605 2025-08-19 cs.CV 57%

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

Wenjie Liao, Jieyu Yuan, Yifang Xu, Chunle Guo, Zilong Zhang, Jihong Li, Jiachen Fu, Haotian Fan, Tao Li, Junhui Cui, Chongyi Li

机构 * Media Evaluation Lab, ByteDance Inc.(字节跳动公司媒体评估实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12472 2025-08-19 cs.AI 57%

GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?

Yifang Tian, Yaming Liu, Zichun Chong, Zihang Huang, Hans-Arno Jacobsen

机构 * University of Toronto(多伦多大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12213 2025-08-19 eess.SP cs.AI cs.LG 57%

Towards Generalizable Human Activity Recognition: A Survey

Yize Cai, Baoshen Guo, Flora Salim, Zhiqing Hong

机构 * Hong Kong University of Science \& Technology (Guangzhou) China SMART, Massachusetts Institute of Technology Singapore University of New South Wales (UNSW) Australia Hong Kong University of Science \& Technology (Guangzhou) SMART, Massachusetts Institute of Technology University of New South Wales (UNSW)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12212 2025-08-19 cs.LG cs.AI q-bio.QM 57%

ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression

Chuanliu Fan, Zicheng Ma, Jun Gao, Nan Yu, Jun Zhang, Ziqiang Cao, Yi Qin Gao, Guohong Fu

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11709 2025-08-19 cs.CY cs.AI 57%

Navigating the New Landscape: A Conceptual Model for Project-Based Assessment (PBA) in the Age of GenAI

Rajan Kadel, Samar Shailendra, Urvashi Rahul Saxena

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Journal ref World Engineering Education Forum - Global Engineering Deans Council (WEEF-GEDC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12502 2025-08-19 math.LO 50%

A Logic of Stability: Formalizing Similarity in Counterfactual Reasoning

Marta Esteves

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 4 篇

2508.11723 2025-08-19 cs.LG 78%

From Heuristics to Data: Quantifying Site Planning Layout Indicators with Deep Learning and Multi-Modal Data

Qian Cao, Jielin Chen, Junchao Zhao, Rudi Stouffs

机构 * National University of Singapore (NUS)(新加坡国立大学) The Cambridge Centre for Advanced Research and Education in Singapore (CARES)(新加坡剑桥高级研究与教育中心) China Southwest Architectural Design and Research Institute Co., Ltd. (CSWADI)(中国西南建筑规划设计研究院有限公司)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

Comments 42 pages, 32 figures, submitted to Environment and Planning B: Urban Analytics and City Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13138 2025-08-19 cs.HC 50%

Human Digital Twin: Data, Models, Applications, and Challenges

Rong Pan, Hongyue Sun, Xiaoyu Chen, Giulia Pedrielli, Jiapeng Huang

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07955 2025-08-19 cond-mat.supr-con 50%

Revealing isotropic abundant low-energy excitations in UTe$_2$ through complex microwave surface impedance

Arthur Carlton-Jones, Alonso Suarez, Yun-Suk Eo, Ian M. Hayes, Shanta R. Saha, Johnpierre Paglione, Nicholas P. Butch, Steven M. Anlage

专题命中 多模态Agent :multi-modal(abstract)

Comments 29 pages, 23 figures

Journal ref Phys. Rev. B 112, 014519 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11781 2025-08-19 cs.HC 50%

Behavioral and Symbolic Fillers as Delay Mitigation for Embodied Conversational Agents in Virtual Reality

Denmar Mojan Gonzales, Snehanjali Kalamkar, Sophie Jörg, Jens Grubert

专题命中 多模态Agent :multimodal(abstract)

Comments Accepted to IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 10 篇

2508.13072 2025-08-19 cs.AI 83%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12917 2025-08-19 cs.CV 83%

CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction

Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo, Jie Yang, Yuming Fang, Wei Liu

机构 * School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所,上海交通大学) School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(计算机与人工智能学院,江西财经大学) School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition & Institute of Medical Robotics, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所及医学机器人研究所,上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The Paper is Accepted by TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02133 2025-08-19 cs.HC 82%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 82%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11319 2025-08-19 cs.CV cs.AI 81%

GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation

Rafi Ibn Sultan, Chengyin Li, Hui Zhu, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu

机构 * Department of Computer Science, Wayne State University, Detroit, MI, USA 48202(计算机科学系,韦恩州立大学) Department of Radiation Oncology, Henry Ford Health, Detroit, MI, USA 48202(放射肿瘤科,亨利福特健康) Department of Electrical and Computer Engineering, The Ohio State University, Columbus, Ohio, USA 43210(电气与计算机工程系,俄亥俄州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11737 2025-08-19 cs.CV cs.AI cs.CL cs.LG 67%

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia, Yuwei Hu, Shanshan Zhao, Yanqing Ma, Zhichao Wei, Yinglun Li, Lunhao Duan, Jianshan Zhao, Yuxuan Han, Haijun Li, Wanying Chen, Junke Tang, Chengkun Hou, Zhixing Du, Tianli Zhou, Wenjie Zhang, Huping Ding, Jiahe Li, Wen Li, Gui Hu, Yiliang Gu, Siran Yang, Jiamang Wang, Hailong Sun, Yibo Wang, Hui Sun, Jinlong Huang, Yuping He, Shengze Shi, Weihong Zhang, Guodong Zheng, Junpeng Jiang, Sensen Gao, Yi-Feng Wu, Sijia Chen, Yuhui Chen, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12036 2025-08-19 cs.CV cs.AI 62%

Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq

机构 * Amity Centre for Artificial Intelligence, Amity University, Noida(阿米蒂人工智能中心,阿米蒂大学,诺伊达)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 4 figures Submitted to AAAI 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12609 2025-08-19 cs.CV 57%

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Lexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu, Hongxiang Li, Lu Ma, Jinghan Ru, Yuexian Zou

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01212 2025-08-19 cs.CV cs.HC 57%

Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference

Yitong Zhu, Zhuowen Liang, Yiming Wu, Tangyao Li, Yuyang Wang

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科学与技术大学(广州)) Nanyang Technological University(南洋理工大学) Nanyang Technological University Singapore(南洋理工大学新加坡)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12418 2025-08-19 cs.LG 50%

Bi-Axial Transformers: Addressing the Increasing Complexity of EHR Classification

Rachael DeVries, Casper Christensen, Marie Lisandra Zepeda Mendoza, Ole Winther

机构 * University of Copenhagen(哥本哈根大学) Novo Nordisk A/S(诺华北欧制药有限公司) Novo Nordisk Research Center Oxford Ltd.(诺华奥克斯研究中心) Technical University of Denmark(丹麦技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 18 pages, 7 figures. Submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他多模态 4 篇

2505.12861 2025-08-19 cs.CV 83%

RMMSS: Towards Advanced Robust Multi-Modal Semantic Segmentation with Hybrid Prototype Distillation and Feature Selection

Jiaqi Tan, Xu Zheng, Yang Liu

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15747 2025-08-19 cs.LG cs.AI 83%

Multi-modal Integration Analysis of Alzheimer's Disease Using Large Language Models and Knowledge Graphs

Kanan Kiguchi, Yunhao Tu, Katsuhiro Ajito, Fady Alnajjar, Kazuyuki Murase

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 38 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏