arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-04 至 2025-09-04 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2509.02915 2025-09-04 cs.CL 83%

English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM

Taekyung Ahn, Hosung Nam

机构 * Enuma, Inc.(Enuma公司) Korea University(韩国大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05779 2025-09-04 cs.CV 79%

A 3D Multimodal Feature for Infrastructure Anomaly Detection

Yixiong Jing, Wei Lin, Brian Sheil, Sinan Acikgoz

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学) Department of Geotechnical Engineering, College of Civil Engineering, Tongji University(地质工程系,同济大学土木学院) Construction Engineering, University of Cambridge(建设工程,剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15867 2025-09-04 cs.CV cs.AI 79%

TruthLens: Visual Grounding for Universal DeepFake Reasoning

Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03385 2025-09-04 cs.CV 70%

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

Reina Ishikawa, Ryo Fujii, Hideo Saito, Ryo Hachiuma

机构 * Keio University(庆应大学) NVIDIA(英伟达)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 70%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16423 2025-09-04 cs.CV cs.LG 70%

GAEA: A Geolocation Aware Conversational Assistant

Ron Campos, Ashmal Vayani, Parth Parag Kulkarni, Rohit Gupta, Aizan Zafar, Aritra Dutta, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments The dataset and code used in this submission is available at: https://ucf-crcv.github.io/GAEA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13037 2025-09-04 eess.IV cs.AI cs.CV 62%

Towards Cardiac MRI Foundation Models: Comprehensive Visual-Tabular Representations for Whole-Heart Assessment and Beyond

Yundi Zhang, Paul Hager, Che Liu, Suprosanna Shit, Chen Chen, Daniel Rueckert, Jiazhen Pan

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02708 2025-09-04 cs.LG cs.AI cs.CL 62%

Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios

Yunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou, Yanggan Gu, Jungang Li, Jingyu Wang, Peijie Jiang, Aiwei Liu, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学) Ant Group(蚂蚁集团) Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00045 2025-09-04 cs.CV cs.PF 57%

Performance is not All You Need: Sustainability Considerations for Algorithms

Xiang Li, Chong Zhang, Hongpeng Wang, Shreyank Narayana Gowda, Yushi Li, Xiaobo Jin

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) The Chinese University of Hong Kong(香港中文大学) University of Nottingham(诺丁汉大学) University of Sydney(悉尼大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 18 pages, 6 figures. Accepted Chinese Conference on Pattern Recognition and Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04233 2025-09-04 eess.IV cs.CV 57%

Grid-Reg: Detector-Free Gridized Feature Learning and Matching for Large-Scale SAR-Optical Image Registration

Xiaochen Wei, Weiwei Guo, Zenghui Zhang, Wenxian Yu

机构 * Shanghai Key Laboratory of Intelligent Sensing and Recognition(上海智能感知与识别重点实验室) Shanghai Jiao Tong University(上海交通大学) Center of Digital Innovation(数字创新中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏