arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-15 至 2025-12-15 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 8 篇

2507.13092 2025-12-15 cs.LG cs.HC 88%

Uncertainty-Aware Cross-Modal Knowledge Distillation with Prototype Learning for Multimodal Brain-Computer Interfaces

具有原型学习的不确定性感知跨模态知识蒸馏用于多模态脑机接口

Hyo-Jeong Jang, Hye-Bin Shin, Seong-Whan Lee

机构 * Department of Brain and Cognitive Engineering, Korea University(脑科学与认知工程系,韩国大学) Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract)

AI总结 本文提出一种具有原型学习的跨模态知识蒸馏框架,通过缓解模态和标签不一致问题,提升多模态脑机接口中EEG的分类和回归性能。

Comments Accepted to SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11558 2025-12-15 cs.CV cs.AI cs.CL 85%

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

DentalGPT: 促进牙科多模态复杂推理的激励机制

Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng, Yanchao Li, Jingyi Liang, Junying Chen, Yunjin Yang, Jiajun You, Shuzhi Deng, Tongfei Wang, Wanting Chen, Chunxiu Hao, Ruiqi Xie, Zhenwei Wen, Xiangyi Feng, Zou Ting, Jin Zou Lin, Jianquan Li, Guangjun Yu, Liangyi Chen, Junwen Wang, Shan Jiang, Benyou Wang

机构 * Shenzhen Stomatology Hospital (Pingshan) of Southern Medical University(南方医科大学深圳口腔医院(平山)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) State Key Laboratory of Membrane Biology, Beijing Key Laboratory of Cardiometabolic Molecular Medicine, Institute of Molecular Medicine, National Biomedical Imaging Center, School of Future Technology, Peking University(北京大学膜生物学国家重点实验室、北京心代谢分子医学重点实验室、分子医学研究院、国家生物医学成像中心、未来技术学院) Freedom AI Division of Applied Oral Sciences & Community Dental Care Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院应用口腔科学与社区牙科护理系) Beijing Institute of Collaborative Innovation(北京协同创新研究院) National Health Data Institute, Shenzhen(深圳国家健康数据研究院) Shenzhen Loop Area Institute(深圳河套学院) Shenzhen Institute of Big Data(深圳大数据研究院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DentalGPT通过高质量数据和强化学习提升牙科多模态推理能力,实现优于现有模型的诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15351 2025-12-15 cs.AI cs.CV 81%

Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration

Octopus: 六种能力协同的代理多模态推理

Yifu Guo, Zishan Xu, Zhiyuan Yao, Yuquan Lu, Jiaye Lin, Sen Hu, Zhenheng Tang, Huacan Wang, Ronghao Chen

机构 * Sun Yat-sen University(中山大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Tsinghua University(清华大学) Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 Octopus通过六种能力协同实现多模态代理推理,有效提升复杂任务中的自主探索与动态能力选择能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11215 2025-12-15 cs.CV 79%

SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection

SmokeBench: 评估多模态大语言模型用于野火烟雾检测

Tianye Qi, Weihao Li, Nick Barnes

机构 * Australian National University(澳大利亚国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 SmokeBench评估多模态大语言模型在野火烟雾检测中的性能,发现模型在烟雾定位方面存在显著局限,尤其在早期阶段。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11074 2025-12-15 cs.CL cs.AI cs.LG cs.MM 67%

MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data

MultiScript30k:利用多语言嵌入扩展跨脚本平行数据

Christopher Driggers-Ellis, Detravious Brinkley, Ray Chen, Aashish Dhawan, Daisy Zhe Wang, Christan Grant

机构 * Computer and Information Science and Engineering(计算机与信息科学与工程)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 MultiScript30k通过多语言嵌入扩展Multi30k数据集,支持多种脚本的全球语言,提升多模态机器翻译的多样性与覆盖范围。

Comments 7 pages, 2 figures, 5 tables. Not published at any conference at this time

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11567 2025-12-15 cs.CL cs.MM 62%

Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet

扩展议会语料库以包含议员的推文:利用MultiParTweet进行自动标注和评估

Mevlüt Bagci, Ali Abusaleh, Daniel Baumartz, Giueseppe Abrami, Maxim Konca, Alexander Mehler

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.MM

AI总结 本文提出MultiParTweet,通过自动标注和人工验证,整合了多语言推文及媒体内容,展示了模型间互为预测性及多模态注释的优越性。

Comments Submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11323 2025-12-15 cs.AI 57%

CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving

CAPTURE:一个用于LVLM在CAPTCHA解决中的基准和评估

Jianyi Zhang, Ziyin Zhou, Xu Ji, Shizhao Liu, Zhangchi Zhao

机构 * Beijing Electric Science and Technology Institute(北京电子科技研究所)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

AI总结 CAPTURE是一个专为LVLM设计的CAPTCHA基准,涵盖4种主要类型和25种子类型,旨在全面评估LVLM在解决CAPTCHA任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10257 2025-12-15 cs.HC 50%

Reject or Not?: A Benchmark for Voice Assistant Query Rejection in Smart Home Scenario and an Improved Method Based on LLMs

拒绝或不?:智能家庭场景下的语音助手查询拒绝基准及基于LLM的改进方法

Huichao Men, Yizhen Hu, Yingyang He, Yu Gao, Xiaofeng Mou, Yi Xu

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出面向智能家庭场景的语音助手查询拒绝基准及基于LLM的改进方法,通过构建多模态数据集和三级协作架构提升拒绝准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏