arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-01 至 2025-12-01 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 18 篇

2511.22055 2025-12-01 cs.CV cs.MM 84%

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

OralGPT-Omni: 一种多功能的牙科多模态大语言模型

Jing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan, Wenkai Zhou, Kaixin Guo, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Yanqi Yang, Qiankun Li, Hao Tang, James Kit-Hon Tsoi, Linlin Shen, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) School of Biomedical Engineering, Southern Medical University(南方医科大学生物医学工程学院) Singapore University of Technology and Design(新加坡科技与设计大学) University of Auckland(奥克兰大学) University of Science and Technology of China(中国科学技术大学) School of Computer Science, Peking University(北京大学计算机学院) College of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

AI总结 OralGPT-Omni是一种专门用于牙科的多模态大语言模型,通过TRACE-CoT数据集和四阶段训练范式,实现了对牙科图像的高效分析和高准确率评估。

Comments 47 pages, 42 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22404 2025-12-01 cs.CV 83%

UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data

UAV-MM3D: 一种大规模合成基准,用于多模态数据下的无人机三维感知

Longkun Zou, Jiale Wang, Rongqin Liang, Hai Wu, Ke Chen, Yaowei Wang

机构 * Pengcheng Laboratory(鹏城实验室) University of Southern California(南加州大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UAV-MM3D通过多模态合成数据提升无人机三维感知能力,提供高保真数据集和多任务基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07984 2025-12-01 cs.CV 83%

SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing

SAMChat:引入链式推理和GRPO以增强小规模遥感遥感小语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中东部技术大学(METU)电子与电气工程系)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SAMChat通过引入链式推理和GRPO,专为遥感影像分析优化,实现了在开放描述和分类任务上的高精度表现。

Comments Accepted to Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) Special Issue on Foundation and Large Vision Models for Remote Sensing. Code and dataset are available at https://github.com/aybora/SAMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12861 2025-12-01 cs.CL cs.CV 81%

From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models

从感知到推理:深度思考赋能多模态大语言模型

Wenxin Zhu, Andong Chen, Yuchen Song, Kehai Chen, Conghui Zhu, Ziyan Chen, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Global Tone Communication Technology(全球语音通信技术)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出多模态链式思维方法,旨在提升多模态大语言模型的推理能力,通过系统性综述分析其理论基础、实现方法及未来发展方向。

Comments Survey; 7 figures, 3 tables, 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23269 2025-12-01 cs.AI 79%

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

OctoMed:面向尖端多模态医疗推理的数据配方

Timothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin, Reuben Tan, Tristan Naumann, Junjie Hu, Hoifung Poon

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 OctoMed通过结构化推理轨迹的数据配方,提升医疗多模态推理模型的性能和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08303 2025-12-01 cs.CV 79%

Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers

通过多模态视觉序列变压器推进语义未来预测

Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

机构 * Archimedes, Athena Research Center(阿基米德研究中心) National Technical University of Athens(希腊国家技术大学) University of Crete(克里特大学) IACM-Forth(第四研究机构(IACM-Forth))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 FUTURIST通过多模态视觉序列变压器架构实现高效的多模态未来语义预测,提升预测精度并简化训练流程。

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11414 2025-12-01 cs.CL 79%

Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization

细粒度且可解释的事实性评估用于多模态摘要

Yue Zhang, Jingxuan Zuo, Ke Su, Liqiang Jing

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出两种细粒度且可解释的评估框架,用于评估多模态摘要模型的事实性,适用于不同应用场景,并通过实验验证了其有效性。

Comments project link: https://github.com/for4WARD/FaithfulnessEvaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21795 2025-12-01 cs.CR 78%

Advanced Data Collection Techniques in Cloud Security: A Multi-Modal Deep Learning Autoencoder Approach

云安全中的高级数据收集技术:一种多模态深度学习自编码器方法

Aamiruddin Syed, Mohammed Ilyas Ahmad

专题命中 多模态评测 :multi-modal(title,abstract)

AI总结 本文提出一种多模态深度学习自编码器方法,通过整合多种数据源和模态,提升云安全中的异常检测与分类性能。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00369 2025-12-01 cs.CV 74%

Automated segmentation of pediatric neuroblastoma on multi-modal MRI: Results of the SPPIN challenge at MICCAI 2023

多模态MRI上儿科神经母细胞瘤自动分割:2023年MICCAI SPPIN挑战赛结果

M. A. D. Buser, D. C. Simons, M. Fitski, M. H. W. A. Wijnen, A. S. Littooij, A. H. ter Brugge, I. N. Vos, M. H. A. Janse, M. de Boer, R. ter Maat, J. Sato, S. Kido, S. Kondo, S. Kasai, M. Wodzinski, H. Muller, J. Ye, J. He, Y. Kirchhoff, M. R. Rokkus, G. Haokai, S. Zitong, M. Fernández Patón, D. Veiga-Canuto, D. G. Ellis, M. R. Aizenberg, B. H. M. van der Velden, H. Kuijf, A. De Luca, A. F. W. van der Steeg

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

AI总结 SPPIN挑战赛通过多模态MRI自动分割神经母细胞瘤,展示了预训练网络在小数据集中的有效性,但小肿瘤分割仍需改进。

Comments 23 pages, 6 figures

Journal ref Bioengineering, 12(11), 1157 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18842 2025-12-01 cs.CV 74%

Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset

通过大规模多模态数据集增强描述性图像质量评估

Zhiyuan You, Jinjin Gu, Xin Cai, Zheyuan Li, Kaiwen Zhu, Chao Dong, Tianfan Xue

机构 * The Chinese University of Hong Kong(香港中文大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Sofia University(索菲亚大学) University of Macau(澳门大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

AI总结 本研究提出DepictQA-Wild模型,通过构建大规模多模态数据集提升图像质量评估的准确性和实用性。

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11561 2025-12-01 cs.CV 70%

Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution

利用大规模语言模型回归准确的图像质量评分使用分数分布

Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, Chao Dong

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Multimedia Laboratory, The Chinese University of Hong Kong(香港中文大学多媒体实验室) Shanghai AI Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学) CPII under InnoHK(创新香港下的CPII)

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本研究提出基于分布的DeQA-Score模型,通过离散化评分分布为软标签,提升图像质量评分的准确性和一致性。

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21735 2025-12-01 cs.CL cs.AI cs.CV 67%

Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting

弥合AI与放射科医生在胸部X光报告中的性能差距

Harshita Sharma, Maxwell C. Reynolds, Valentina Salvatelli, Anne-Marie G. Sykes, Kelly K. Horst, Anton Schwaighofer, Maximilian Ilse, Olesya Melnichenko, Sam Bond-Taylor, Fernando Pérez-García, Vamshi K. Mugu, Alex Chan, Ceylan Colak, Shelby A. Swartz, Motassem B. Nashawaty, Austin J. Gonzalez, Heather A. Ouellette, Selnur B. Erdal, Beth A. Schueler, Maria T. Wetscherek, Noel Codella, Mohit Jain, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Stephanie Hyland, Panos Korfiatis, Ashish Khandelwal, Javier Alvarez-Valle

机构 * Microsoft(微软公司) Mayo Clinic(梅奥诊所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MAIRA-X通过多模态AI模型在胸部X光报告生成中提升词汇质量、临床正确性和L&T准确性,有效辅助放射科医生,尤其在高患者量的临床环境中

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22178 2025-12-01 cs.CV cs.AI 62%

Enhanced Graph Convolutional Network with Chebyshev Spectral Graph and Graph Attention for Autism Spectrum Disorder Classification

增强型图卷积网络结合切比雪夫谱图与图注意力用于自闭症谱系障碍分类

Adnan Ferdous Ashrafi, Hasanul Kabir

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Islamic University of Technology(伊斯兰科技大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出结合切比雪夫谱图卷积和图注意力网络的增强型图卷积网络,用于提高自闭症谱系障碍分类的准确性。

Comments 6 pages, 2 figures, 2 tables, Accepted and presented at Image and Vision Computing New Zealand (IVCNZ) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23214 2025-12-01 cs.CV 57%

Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation

面向半受控工业环境的零样本多准则视觉质量检测:通过实时3D数字孪生模拟

Jose Moises Araya-Martinez, Gautham Mohan, Kenichi Hayakawa Bolaños, Roberto Mendieta, Sarvenaz Sardari, Jens Lambrecht, Jörg Krüger

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于实时3D数字孪生模拟的零样本多准则视觉质量检测框架,用于半受控工业环境中的高效质量检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22576 2025-12-01 cs.MM 57%

A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization

一种用于故事可视化多元文化分析的渐进评估框架

Janak Kapuriya, Ali Hatami, Paul Buitelaar

专题命中 多模态评测 :MLLM(abstract);分类 cs.MM

AI总结 本文提出一种渐进多元文化评估框架,通过五个新指标评估故事可视化模型在不同文化下的表现,揭示模型在文化适当性和视觉美学上的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21716 2025-12-01 cs.CL 57%

An Optimized Machine Learning Classifier for Detecting Fake Reviews Using Extracted Features

一种用于通过提取特征检测虚假评论的优化机器学习分类器

Shabbir Anees, Anshuman, Ayush Chaurasia, Prathmesh Bogar

机构 * Indian Institute of Information Technology Vadodara(印度瓦达拉信息科技大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

AI总结 本文提出了一种结合HHO优化和堆叠集成的机器学习方法,用于高精度识别由人工智能生成的虚假评论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06597 2025-12-01 cs.RO 50%

LiHRA: A LiDAR-Based HRI Dataset for Automated Risk Monitoring Methods

LiHRA:基于LiDAR的人机交互风险监测数据集

Frederik Plahl, Georgios Katranis, Ilshat Mamaev, Andrey Morozov

机构 * Proximity Robotics & Automation GmbH(近距机器人与自动化有限公司) Institute of Industrial Automation and Software Engineering, University of Stuttgart(工业自动化与软件工程学院,斯图加特大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 LiHRA数据集通过多模态数据支持人机交互风险监测方法的开发,提供高分辨率LiDAR数据和真实碰撞事件,用于训练和评估RM算法。

Comments Preprint of final paper that will appear in the Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09042 2025-12-01 eess.SY cs.LG cs.SY 50%

MAKO: Meta-Adaptive Koopman Operators for Learning-based Model Predictive Control of Parametrically Uncertain Nonlinear Systems

MAKO:基于元学习的Koopman算子用于参数不确定非线性系统的基于学习的模型预测控制

Minghao Han, Kiwan Wong, Adrian Wing-Keung Law, Xunyuan Yin

机构 * Water Research Institute (NEWRI), Nanyang Technological University, Singapore(新跃大学水研究 institute(NEWRI)) School of Chemistry, Chemical Engineering and Biotechnology, Nanyang Technological University, Singapore(新跃大学化学、化工与生物技术学院) Soft Robotics Lab, ETH Zurich, Switzerland(苏黎世联邦理工学院软机器人实验室) Department of Civil and Environmental Engineering, National University of Singapore, Singapore(新加坡国立大学土木与环境工程系)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 MAKO通过元学习方法实现对参数不确定非线性系统的高效建模与预测控制,提升系统稳定性和控制效果。

详情

展开后加载摘要…

URL PDF HTML 收藏