arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-06 至 2026-01-06 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 18 篇

2508.06142 2026-01-06 cs.CV 83%

SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models

SDEval: 多模态大语言模型的安全动态评估

Hanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang, Xiangyang Zhu

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SDEval通过动态调整安全基准分布和复杂度,提升多模态大语言模型的安全评估效果,缓解数据污染并揭示模型安全局限。

Comments AAAI 2026 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06057 2026-01-06 cs.CL cs.MM 81%

MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection

注意你的推理:一种用于多模态立场检测双推理的元认知直观-反思网络

Bingbing Wang, Zhengda Jin, Bin Liang, Wenjie Li, Jing Li, Ruifeng Xu, Min Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

AI总结 MIND通过元认知直观-反思机制,在多模态立场检测中实现双推理,提升立场判断的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00907 2026-01-06 eess.IV cs.AI cs.CV cs.LG 81%

Placenta Accreta Spectrum Detection using Multimodal Deep Learning

利用多模态深度学习进行胎盘增生谱检测

Sumaiya Ali, Areej Alhothali, Sameera Albasri, Ohoud Alzamzami, Ahmed Abduljabbar, Muhammad Alwazzan

机构 * Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University(计算机科学系,计算与信息科技学院,国王阿卜杜勒阿齐兹大学) Department of Radiology, King Abdulaziz University Hospital(放射科,国王阿卜杜勒阿齐兹大学医院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出多模态深度学习方法,通过整合3D MRI和2D US数据,提高胎盘增生谱检测的准确率和AUC值,有效提升产前风险评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22535 2026-01-06 cs.AI cs.CL 81%

OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models

OFFSIDE: 多模态大语言模型中误信信息消除的基准测试

Hao Zheng, Zirui Pang, Ling li, Zhijie Deng, Yuhan Pu, Zhaowei Zhu, Xiaobo Xia, Jiaheng Wei

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) BIAI, ZJUT & D5Data.ai(BIAI、浙江工业大学及D5Data.ai) National University of Singapore(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 OFFSIDE通过足球转会谣言数据集,评估多模态大语言模型中误信信息消除的挑战与方法,揭示现有技术的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03663 2026-01-06 cs.CL cs.CV 81%

UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG

UNIDOC-BENCH: 一个统一的文档中心多模态RAG基准

Xiangyu Peng, Can Qin, Zeyuan Chen, Ran Xu, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 UniDoc-Bench是首个大规模文档中心多模态RAG基准,通过多模态问答对评估文本-图像融合与联合检索性能,揭示多模态嵌入不足及视觉上下文补充机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16505 2026-01-06 cs.CV cs.MM 81%

TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning

TraveLLaMA: 一种基于大规模数据集和结构化推理的多模态旅行助手

Meng Chu, Yukang Chen, Haokun Gui, Shaozuo Yu, Yi Wang, Jiaya Jia

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 TraveLLaMA通过结构化推理框架和大规模数据集,提升旅行推荐和场景理解能力,实现用户满意度和系统性能的显著提升。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01033 2026-01-06 eess.SP 78%

System-Level Comparison of Multimodal and In-Band mmWave Sensing for Beam Prediction in 6G ISAC

6G ISAC中多模态与带内毫米波感知的系统级比较:用于基站预测

Abidemi Orimogunje, Hyunwoo Park, Igbafe Orikumhi, Sunwoo Kim, Dejan Vukobratovic

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出了一种系统级框架,通过多模态融合评估毫米波与体外传感器在6G ISAC中的波束预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00843 2026-01-06 cs.AI 74%

OmniNeuro: A Multimodal HCI Framework for Explainable BCI Feedback via Generative AI and Sonification

OmniNeuro:一种通过生成式AI和声音化实现可解释BCI反馈的多模态人机交互框架

Ayda Aghaei Nia

机构 * Institute for Artificial Intelligence Researcher(人工智能研究院)

专题命中 多模态评测 :multimodal(title);分类 cs.AI

AI总结 OmniNeuro通过生成式AI和声音化技术,提供可解释的BCI反馈,提升用户神经可塑性与解码准确性。

Comments 16 pages, 7 figures, 3 tables. Source code and implementation available at: https://github.com/ayda-aghaei/OmniNeuro. Highlights the use of LLMs (Gemini) and Quantum probability formalism for real-time BCI explainability

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01868 2026-01-06 cs.CL 70%

DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs

DermoGPT: 开放权重和开放数据用于基于形态的皮肤病学推理大语言模型

Jinghan Ru, Siyuan Yan, Yuguo Yin, Yuexian Zou, Zongyuan Ge

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Faculty of Information Technology, Monash University(信息科技学院,莫纳什大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 DermoGPT通过开放权重和数据提升皮肤病学推理的MLLM性能,实现与专家诊断流程的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01188 2026-01-06 cs.RO cs.CV 70%

DST-Calib: A Dual-Path, Self-Supervised, Target-Free LiDAR-Camera Extrinsic Calibration Network

DST-Calib: 一种双路径、自监督、无目标的激光雷达-摄像头外参校准网络

Zhiwei Huang, Yanwei Fu, Yi Zhou, Xieyuanli Chen, Qijun Chen, Rui Fan

机构 * Department of Control Science & Engineering, the College of Electronics & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学) Department of Mechanical Engineering, the School of Mechanical Engineering, Tongji University(机械工程系,机械工程学院,同济大学) School of Robotics, Hunan University(机器人学院,湖南大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科学技术大学) Department of Control Science & Engineering, the College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University(控制科学与工程系,电子与信息工程学院,上海智能自主系统研究院,智能自主系统国家重点实验室,智能自主系统前沿科学中心,同济大学) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 DST-Calib通过双路径自监督方法,无需目标实现激光雷达与摄像头的高效外参校准,提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03104 2026-01-06 cs.AI 70%

ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning

通过合成数据对齐时间序列,利用大语言模型进行增强的理解与推理

Zhe Xie, Zeyan Li, Xiao He, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen, Rui Shi, Dan Pei

机构 * Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 ChatTS通过合成数据生成方法,首次实现了多变量时间序列的多模态大语言模型,显著提升了时间序列的理解与推理性能。

Comments accepted by VLDB' 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02008 2026-01-06 cs.AI cs.CV 62%

XAI-MeD: Explainable Knowledge Guided Neuro-Symbolic Framework for Domain Generalization and Rare Class Detection in Medical Imaging

XAI-MeD: 可解释知识引导的神经符号框架用于医学影像中的领域泛化和稀有类别检测

Midhat Urooj, Ayan Banerjee, Sandeep Gupta

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 XAI-MeD通过整合临床知识的神经符号框架,提升医学影像中领域泛化和稀有类别检测的性能,实现更鲁棒和可解释的多模态医学AI。

Comments Accepted at AAAI Bridge Program 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23413 2026-01-06 cs.CV 57%

Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment

弥合认知鸿沟:面向艺术图像美学评估的层次描述学习

Henglin Liu, Nisha Huang, Chang Liu, Jiangpeng Yan, Huijuan Huang, Jixuan Ying, Tong-Yee Lee, Pengfei Wan, Xiangyang Ji

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出RAD数据集和ArtQuant框架,通过联合描述生成和LLM解码器提升艺术图像美学评估的准确性和效率。

Comments AAAI2026,Project Page:https://github.com/Henglin-Liu/ArtQuant

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18829 2026-01-06 cs.AI 57%

HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare

HARBOR:面向行为医疗的综合自适应风险评估模型

Aditya Siddhant

机构 * Aditya Siddhant(独立研究者)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 HARBOR模型通过整合多模态数据,有效预测行为医疗中的情绪和风险评分,实现69%的准确率,优于传统方法和现有LLMs。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23252 2026-01-06 cs.CV 57%

DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection

DGE-YOLO:双分支聚集与注意力用于准确的无人机目标检测

Kunwei Lv, Zhiren Xiao, Hang Ren, Ping Lan

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 DGE-YOLO 通过双分支架构和高效多尺度注意力机制,提升多模态无人机目标检测的准确性和鲁棒性。

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00839 2026-01-06 cs.CV 57%

Unified Review and Benchmark of Deep Segmentation Architectures for Cardiac Ultrasound on CAMUS

针对心脏超声的深度分割架构的统一回顾与基准测试:CAMUS

Zahid Ullah, Muhammad Hilal, Eunsoo Lee, Dragan Pamucar, Jihie Kim

机构 * Department of Computer Science Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea Department of Semiconductor System Engineering, Sejong University, Seoul, 05006, Republic of Korea Department of Operations Research Statistics, Faculty of Organizational Sciences, University of Belgrade, Belgrade, Serbia Department of Industrial Engineering \& Management, Yuan Ze University, Taoyuan City 320315, Taiwan Department of Applied Mathematical Science, College of Science Technology, Korea University, Sejong 30019, Republic of Korea

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文通过统一基准测试评估了三种心脏超声分割架构的性能,提出标准化预处理方法,并展望了自监督和多模态标注技术的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24029 2026-01-06 cs.RO cs.HC 50%

Evaluation of Impression Difference of a Domestic Mobile Manipulator with Autonomous and/or Remote Control in Fetch-and-Carry Tasks

对具有自主和/或远程控制的国产移动机械臂在抓取和搬运任务中印象差异的评估

Takashi Yamamoto, Hiroaki Yaguchi, Shohei Kato, Hiroyuki Okada

专题命中 多模态评测 :multimodal(abstract)

AI总结 本研究评估了国产移动机械臂在抓取和搬运任务中自主与远程控制模式对用户印象的影响。

Comments Published in Advanced Robotics (2020). v2 updates Abstract/Comments (metadata only); paper content unchanged. Please cite: Advanced Robotics 34(20):1291-1308, 2020. https://doi.org/10.1080/01691864.2020.1780152

Journal ref Advanced Robotics, 34(20):1291-1308, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00851 2026-01-06 physics.ins-det cond-mat.mtrl-sci cs.LG 50%

Autonomous battery research: Principles of heuristic operando experimentation

自主电池研究:启发式在位实验原理

Emily Lu, Gabriel Perez, Peter Baker, Daniel Irving, Santosh Kumar, Veronica Celorrio, Sylvia Britto, Thomas F. Headen, Miguel Gomez-Gonzalez, Connor Wright, Calum Green, Robert Scott Young, Oleg Kirichek, Ali Mortazavi, Sarah Day, Isabel Antony, Zoe Wright, Thomas Wood, Tim Snow, Jeyan Thiyagalingam, Paul Quinn, Martin Owen Jones, William David, James Le Houx

机构 * ISIS Neutron & Muon Source, Rutherford Appleton Laboratory(ISIS中子与穆子源、拉瑟福德-苹果顿实验室) The Faraday Institution(法拉第机构) Diamond Light Source, Rutherford Appleton Laboratory(Diamond光源、拉瑟福德-苹果顿实验室) University of Cambridge, The Old Schools, Trinity Ln(剑桥大学、旧校舍、三一街) Imperial College London, Department of Mechanical Engineering(伦敦帝国理工学院、机械工程系)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文提出启发式在位实验框架,利用AI和数字孪生技术主动捕捉电池退化中的罕见事件,提升实验效率和数据可靠性。

Comments 38 pages, 14 figures. Includes a detailed technical review of the POLARIS, BAM, DRIX, M-Series, and B18 electrochemical cells in the Supplementary Information

详情

展开后加载摘要…

URL PDF HTML 收藏