arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9150 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9150 篇

1608.02289 2016-08-09 cs.CV cs.CL cs.MM 82%

Detecting Sarcasm in Multimodal Social Platforms

Rossano Schifanella, Paloma de Juan, Joel Tetreault, Liangliang Cao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 10 pages, 3 figures, final version published in the Proceedings of ACM Multimedia 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08140 2026-04-10 cs.CR cs.AI cs.MM cs.NI 82%

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

基于LLM的多模态推理用于加密流量解释:一个基准

Longgang Zhang, Xiaowei Fu, Fuxiang Huang, Lei Zhang

机构 * School of Microelectronics and Communication Engineering, Chongqing University(重庆大学微电子与通信工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 本文提出BGTD基准和mmTraffic框架,通过结合原始字节与结构化注释,实现可解释的加密流量解释,生成高保真的人可读报告,同时保持高分类准确率。

Comments Project page \url{https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark}

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23499 2026-03-12 cs.CV cs.CL cs.LG 82%

Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional

多模态数据光谱:多模态数据集是多维的

Divyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit Chopra

机构 * New York University(纽约大学) GenentechCIFAR(基因泰克CIFAR) New York University Grossman School of Medicine(纽约大学格罗斯曼医学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL;multimodal(comments)

AI总结 本文通过多模态大语言模型分析23个视觉问答基准,揭示了不同模态在目标任务中的依赖性差异,发现部分基准因设计缺陷放大了图像依赖性。

Comments Accepted to ICLR 2026. Code available at https://github.com/divyam3897/multimodal-spectrum

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17337 2025-09-23 cs.AI cs.CL 82%

LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code

Ala Jararweh, Michael Adams, Avinash Sahu, Abdullah Mueen, Afsah Anwar

机构 * Department of Computer Science, The University of New Mexico(计算机科学系,新墨西哥大学) Comprehensive Cancer Center, The University of New Mexico(综合癌症中心,新墨西哥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref A. Jararweh, M. Adams, A. Sahu, A. Mueen and A. Anwar, "LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code," 2025 5th Intelligent Cybersecurity Conference (ICSC), Tampa, FL, USA, 2025, pp. 232-241

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09401 2025-06-10 cs.CV cs.AI cs.RO 82%

MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

Ruiyuan Lyu, Jingli Lin, Tai Wang, Shuai Yang, Xiaohan Mao, Yilun Chen, Runsen Xu, Haifeng Huang, Chenming Zhu, Dahua Lin, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhiyuan College, Shanghai Jiao Tong University(上海交通大学紫阳学院) CPII under InnoHK(创新香港下的CPII)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Follow-up of EmbodiedScan (camera-ready version). A multi-modal 3D dataset with the most-ever comprehensive language annotations for 3D-LLMs. Project page: https://tai-wang.github.io/mmscan/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10483 2025-05-16 cs.CV cs.AI 82%

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

Yi Li, Haonan Wang, Qixiang Zhang, Boyu Xiao, Chenchang Hu, Hualiang Wang, Xiaomeng Li

机构 * David S. Hippocampus Department of Computer Science(戴维·S·海马科斯部门计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments UniEval is the first evaluation framework designed for unified multimodal models, including a holistic benchmark UniBench and the UniScore metric

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01523 2025-03-03 cs.CL cs.AI 82%

GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Hongzhan Lin, Ziyang Luo, Bo Wang, Ruichao Yang, Jing Ma

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The first work to benchmark Large Multimodal Models in safety insight on social media

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14191 2025-02-21 cs.CV cs.AI 82%

Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Michihiro Yasunaga, Luke Zettlemoyer, Marjan Ghazvininejad

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Dataset available at https://github.com/facebookresearch/multimodal_rewardbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09994 2025-01-20 cs.CV cs.AI eess.IV 82%

Multi-Modal Attention Networks for Enhanced Segmentation and Depth Estimation of Subsurface Defects in Pulse Thermography

Mohammed Salah, Naoufel Werghi, Davor Svetinovic, Yusra Abdulrahman

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Pulse thermography, infrared thermography, defect segmentation, multi-modal networks, attention mechanism

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10822 2022-08-24 cs.CV cs.AI cs.HC 82%

Multimodal Across Domains Gaze Target Detection

Francesco Tonini, Cigdem Beyan, Elisa Ricci

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to 24th ACM International Conference on Multimodal Interaction (ICMI 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12352 2021-06-18 cs.CV cs.CL 82%

Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks

Letitia Parcalabescu, Albert Gatt, Anette Frank, Iacer Calixto

专题命中 多模态评测 :cross-modal(title);multimodal(abstract,journal_ref);分类 cs.CV、cs.CL

Comments Paper accepted for publication at MMSR 2021; 13 pages, 3 figures, 7 Tables

Journal ref Proceedings of the 1st Workshop on Multimodal Semantic Representations (MMSR), 2021, Groningen, Netherlands (Online), Association for Computational Linguistics, p. 32--44

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.08267 2020-08-04 cs.CL cs.LG cs.SD eess.AS 82%

Multilogue-Net: A Context Aware RNN for Multi-modal Emotion Detection and Sentiment Analysis in Conversation

Aman Shenoy, Ashish Sardana

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、eess.AS;multimodal(comments)

Comments 10 pages, 3 figures, 5 tables; Published in Proceedings of the Second Grand Challenge and Workshop on Multimodal Language (Challenge-HML) in the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020)

Journal ref Challenge-HML, ACL 2020, 19-28

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27278 2026-08-04 cs.CV 版本更新 81%

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

OVEarth-Bench:面向开放词汇地球观测的类别广度与查询多样性评估

Kaiyu Li, Zepeng Xin, Zixuan Jiang, Jing Fu, Lanxuan Xue, Lingyu Zhang, Xiangyong Cao

专题命中 多模态评测 :MLLM(summary_cn,abstract);分类 cs.CV

AI总结 本文提出OVEarth-Bench基准,从类别广度与查询多样性两方面扩展开放词汇地球观测评估,经实验发现MLLM类方法性能最优,为该领域未来研究提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21400 2026-07-24 cs.RO cs.AI 新提交 81%

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

VoLN:仅视觉的长距离导航——范式、基准和方法

Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu

机构 * Beihang University(北京航空航天大学) Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院)

专题命中 多模态评测 :MLLM(summary_cn,abstract);分类 cs.AI

AI总结 研究提出仅视觉的长距离导航(VoLN)范式,通过VoLN-UAV基准及VoLN-MLLM基线进行实例化。在五个环境测试未见分割中评估,揭示了长距离导航在证据整合、目标匹配和闭环稳定性方面的挑战。

Comments 10 pages, 7 figures, 2 tables. Project page: https://admire-ljb.github.io/VoLN-UAV/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02482 2026-06-30 cs.CV 81%

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

X-Stream: 探索多模态大语言模型作为多流理解的多路复用器

Peiwen Sun, Xudong Lu, Huadai Liu, Yang Bo, Dongming Wu, Huankang Guan, Minghong Cai, Jinpeng Chen, Xintong Guo, Shuhan Li, Fang Liu, Rui Liu, Xiangyu Yue

机构 * MMLab, Chinese University of Hong Kong(中大香港人工智能实验室) Huawei Inc.(华为公司)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multi-modal(abstract);分类 cs.CV

AI总结 为解决多流视频理解评估缺失的问题,提出首个基准X-Stream,包含4220个QA对和932个视频,覆盖多窗口、多视角和多设备场景,并基于信号多路复用理论评估MLLM作为多路复用器的性能,发现现有模型在并发流上仅达约50%分数。

Comments Project Page: https://peiwensun2000.github.io/xstream/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25634 2026-06-25 cs.CV 新提交 81%

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity

SSMNBench: 通过单视图充分性与多视图必要性诊断基于图像的跨视角人-物理解

Tianchen Guo, Chen Liu, Ling Chen, Xin Yu

机构 * The University of Queensland(昆士兰大学) Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所) University of Technology Sydney(悉尼科技大学) Follow Me AI Pty LTD(Follow Me AI有限公司)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出SSMNBench基准,通过单视图充分性(SVS)和多视图必要性(MVN)任务分类,诊断MLLM在跨视角人-物理解中的视觉干扰退化和跨视角融合失败问题。

Comments European Conference on Computer Vision (ECCV). 32 pages, 10 figures. The code is available at: $ \href{https://github.com/gtc-gh/SSMNBench}{\text{SSMNBench}} $

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14958 2026-06-16 cs.CV cs.IR cs.LG 新提交 81%

MVEB: Massive Video Embedding Benchmark

MVEB:大规模视频嵌入基准

Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, Márton Kardos, Niklas Muennighoff, Kenneth Enevoldsen

机构 * Harvard University(哈佛大学) SaluteDevices MIRAI Zendesk Shanghai University of Finance and Economics(上海财经大学) Google LLC Salesforce Cornell University(康奈尔大学) Astera Institute(Astera研究院) Independent Contributor(独立贡献者) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦分校) Barclays(巴克莱银行) Aarhus University(奥胡斯大学) Stanford University(斯坦福大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出MVEB基准,包含23个任务评估33种视频嵌入模型,发现无单一模型占优,音频贡献取决于标注来源,并集成到MTEB生态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03175 2026-06-04 cs.CV cs.RO 81%

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

在值得时询问:面向实例目标导航的成本感知开放式交互

Xunyi Zhao, Sihao Lin, Gengze Zhou, Zerui Li, Shijie Li, Wei Tao, Jiajun Liu, Qi Wu

机构 * Adelaide University(阿德莱德大学) Responsible AI Research Centre, Australian Institute for Machine Learning(负责任人工智能研究中心,澳大利亚机器学习研究所) Institute for Infocomm Research (I2R), A*STAR(信息与通信研究院(I2R),A*STAR) iMotion CSIRO Data61 Project Website(CSIRO Data61项目网站)

专题命中 多模态评测 :MLLM(summary_cn,abstract);分类 cs.CV

AI总结 针对实例目标导航中语言歧义问题,提出一种成本敏感的不确定性减少方法,通过信息增益分析确定有效问题类型,并构建基准测试和加权成功率指标,实现零样本MLLM导航器仅在预期收益大于成本时查询。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15876 2026-05-21 cs.CV 81%

Unlocking Dense Metric Depth Estimation in VLMs

解锁VLMs中的密集度量深度估计

Hanxun Yu, Xuan Qu, Yuxin Wang, Jianke Zhu, Lei Ke

机构 * Zhejiang University(浙江大学) Tencent Hunyuan LLM(腾讯混元大模型) HKUST(香港科技大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 多模态评测 :multimodal(abstract,abstract_cn);multimodal foundation model(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出DepthVLM,一种将单个VLM转换为原生密集几何预测器的简单有效框架,同时保持其多模态能力。通过在LLM主干上附加轻量级深度头,并在统一的视觉-文本监督范式下进行训练,DepthVLM能够在单次前向传递中生成高分辨率深度图和语言输出。此外,还引入了一个统一的室内-室外度量深度基准,实验表明DepthVLM在推理效率、复杂3D空间推理等方面均优于现有VLMs和纯视觉模型。

Comments Project Page: https://depthvlm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10064 2026-04-14 cs.CV 81%

On The Application of Linear Attention in Multimodal Transformers

多模态Transformer中线性注意力的应用

Armin Gerami, Seyedehanita Madani, Ramani Duraiswami

机构 * University of Maryland(马里兰大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV;any-to-any(comments)

AI总结 本文研究了线性注意力在多模态框架中的应用,通过减少计算开销并保持性能,证明了线性注意力在多模态Transformer中的有效性。

Comments Workshop on Any-to-Any Multimodal Learning (Any2Any), CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12820 2026-01-21 cs.CV 81%

A Generalist Foundation Model for Total-body PET/CT Enables Diagnostic Reporting and System-wide Metabolic Profiling

一种通用的基础模型用于全身PET/CT,实现诊断报告和系统层面的代谢分析

Wei Chen, Liang Wu, Shuyi Lu, Yuanyuan Sun, Wenkai Bi, Zilong Yuan, Yaoyao He, Feng Wang, Junchi Ma, Shuyong Liu, Zhaoping Cheng, Xiaoyan Hu, Jianfeng Qiu

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);image-text(abstract);multimodal foundation model(abstract)

AI总结 SDF-HOLO是一种用于全身PET/CT的通用基础模型,通过多模态学习实现诊断报告和系统代谢分析,提升医疗影像处理的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19212 2026-08-21 cs.CL cs.CV 新提交 81%

NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection

NepOOC-M:面向OOC检测的尼泊尔语-英语双语基准及多模态架构的对比分析

Sanjeev Khatiwada

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 针对尼泊尔语OOC虚假信息检测,构建首个尼泊尔语多语言OOC基准NepOOC,评估多模态等架构发现仅文本mBERT性能最优,数据集扩充比架构优化更有效。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18591 2026-08-20 cs.AI cs.CL 新提交 81%

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

轻量级多模态模型能否评估LLM的推理性能?一项面向计算最优文档推理的研究

Zishan Ahmad, Vishal Vaddina

机构 * Phi Labs(菲实验室) Quantiphi(宽梯斐科技)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出轻量级多模态模型DRB,基于新基准BudgetDoc实现文档任务中LLM推理预算的动态分配,在降低成本的同时多数情况下达到或优于固定最大预算的性能,具备跨模型泛化潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18586 2026-08-20 cs.CV cs.AI 新提交 81%

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR:用于评估多模态大语言模型在手写OCR场景中表现的诊断基准

Zinuo Guo, Min Zhang, Bo Jiang

机构 * East China Normal University(华东师范大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究推出OmniHandwritingOCR手写OCR诊断基准,评估13个多模态模型,发现其在复杂公式上表现差且存在幻觉,为诊断模型失效模式提供测试平台。

Comments CIKM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17895 2026-08-19 cs.CL cs.AI 新提交 81%

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

BEAR-Bench:面向多模态模型的双语企业与学术推理基准

Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev

机构 * Yandex Applied AI Institute(Yandex应用人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文针对多模态大语言模型在文本密集专业文档推理能力评估的不足,推出英俄双语的BEAR-Bench基准,评估16个多模态大语言模型并对比幻觉检测方法,发现最强模型仍有明显性能提升空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16233 2026-08-18 eess.IV cs.AI cs.CV 新提交 81%

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

面向不完整与退化前列腺MRI的跨模态生成模型及多中心临床验证

Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Guangyu Wu

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 该研究开发MSCNet跨模态生成框架,可重建前列腺MRI缺失对比、恢复退化扫描,经多中心临床验证,其图像质量符合部分非劣效标准,诊断效能优于基线生成图像,可作为前列腺MRI的辅助技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14389 2026-08-17 cs.CV cs.AI 新提交 81%

GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection

GBU-Palm:用于掌纹呈现攻击检测的多模态视频数据集与基准

Yingjie Ma, Zitong Yu, Wei Jia, Ajay Kumar, Linlin Shen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出GBU-Palm多模态掌纹PAD数据集与基准,通过控制泄漏协议分离身份与攻击谱系,测试4种视频架构,发现环境变化下架构性能下降,RGB-NIR融合未始终优于仅RGB输入,为跨环境掌纹PAD研究提供基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14075 2026-08-17 cs.AI cs.CV cs.DL 新提交 81%

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

通用科学人工智能的实现路径:科学图像的多模态理解

Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden

机构 * TIB Leibniz Information Centre for Science and Technology(莱布尼茨科学与技术信息中心(TIB)) Eindhoven University of Technology(埃因霍温理工大学) Freie Universität Berlin(柏林自由大学) International Center for Chemical and Biological Sciences, University of Karachi(卡拉奇大学国际化学与生物科学中心) National University of Sciences and Technology(国家科技大学) University of Warwick(华威大学) PSG College of Technology(PSG理工学院) Aalto University(阿尔托大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 该研究提出以科学图像多模态理解构建通用科学AI的路径,依托ALD/E-ImageMiner基准与2026年ICDAR竞赛,明确相关任务能力检验维度及长期研究方向,推动可机器执行的科学视觉知识与可验证多模态科学AI发展。

Comments 15 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00056 2026-08-13 cs.CY cs.AI cs.CL 版本更新 81%

How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?

VLMs在帮助人类推断从多模态简答中获得的思维模型质量的有效性如何?

Pritam Sil, Durgaprasad Karnam, Vinay Reddy Venumuddala, Pushpak Bhattacharyya

机构 * Department of Computer Science and Engineering, IIT Bombay(印度理工学院班加罗尔分校计算机科学与工程系) Center for Educational Technology, IIT Bombay(印度理工学院班加罗尔分校教育技术中心) School of Management, Mahindra University(马恒达大学管理学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出MMGrader方法,通过概念图框架从多模态回答推断学生思维模型质量,发现最佳模型在人类水平上表现不足,但能有效辅助教师改进教学。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17332 2026-08-13 cs.CV cs.AI 版本更新 81%

P2MFDS: A Privacy-Preserving Multimodal Fall Detection System for Elderly People in Bathroom Environments

P2MFDS:面向浴室环境老年人的隐私保护多模态跌倒检测系统

Haitian Wang, Yiren Wang, Xinyu Wang, Yumeng Miao, Yuliang Zhang, Yu Zhang, Atif Mansoor

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Department of Computer Science and Software Engineering, The University of Western Australia(西澳大学计算机科学与软件工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对浴室环境老年人跌倒检测的单模态系统准确率受限问题,本文提出P2MFDS多模态系统,融合毫米波雷达与3D振动传感,构建大规模隐私保护数据集,双流网络结合多尺度特征,提升了跌倒检测的准确率与召回率。

Comments Accepted to appear in the 2025 IEEE International Workshop on AIoT and Smart Systems (AIoTSys'25)

详情

展开后加载摘要…

URL PDF HTML 收藏