arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9170 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9170 篇

2308.11020 2023-09-26 cs.CL cs.HC cs.RO 80%

Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors

Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara, Gabriel Skantze

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by 25th ACM International Conference on Multimodal Interaction (ICMI '23), Late-Breaking Results

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05417 2023-07-20 cs.CV 80%

The MONET dataset: Multimodal drone thermal dataset recorded in rural scenarios

Luigi Riz, Andrea Caraffa, Matteo Bortolon, Mohamed Lamine Mekhalfi, Davide Boscaini, André Moura, José Antunes, André Dias, Hugo Silva, Andreas Leonidou, Christos Constantinides, Christos Keleshis, Dante Abate, Fabio Poiesi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Published in Computer Vision and Pattern Recognition (CVPR) Workshops 2023 - 6th Multimodal Learning and Applications Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05542 2021-05-13 cs.CL 80%

!Qué maravilla! Multimodal Sarcasm Detection in Spanish: a Dataset and a Baseline

Khalid Alnajjar, Mika Hämäläinen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to The Third Workshop on Multimodal Artificial Intelligence (MAI-Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.03915 2018-08-07 cs.CL cs.LG stat.ML 80%

Seq2Seq2Sentiment: Multimodal Sequence to Sequence Models for Sentiment Analysis

Hai Pham, Thomas Manzini, Paul Pu Liang, Barnabas Poczos

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 8 pages of content, 11 pages total, 2 figures. Published as a workshop paper at ACL 2018, Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML). 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05502 2026-07-24 cs.CV cs.AI cs.CL 版本更新 80%

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基

Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(东华大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对低资源语言MLLMs文化内涵不足问题,提出MELLA数据集,采用双源策略,结合母语网络图像-替代文本对与生成翻译的图像描述进行监督,经实验表明能减轻文化幻觉,强调数据对齐对低资源语言文化基础多模态理解的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04413 2026-02-05 cs.CL cs.AI cs.MM 80%

History-Guided Iterative Visual Reasoning with Self-Correction

基于历史的迭代视觉推理与自我校正

Xinglong Yang, Zhilin Peng, Zhanzhan Liu, Haochen Shi, Sheng-Jun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 H-GIVR框架通过迭代视觉推理与自我校正,显著提升多模态推理准确性并保持低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17856 2025-05-02 cs.LG eess.SP 80%

Enhancing clinical decision support with physiological waveforms -- a multimodal benchmark in emergency care

Juan Miguel Lopez Alcaraz, Hjalmar Bouma, Nils Strodthoff

专题命中 多模态评测 :multimodal(title,abstract)

Comments Version accepted by Computers in Biology and Medicine: 21 pages, 2 figures, code available under https://github.com/AI4HealthUOL/MDS-ED, dataset available under https://physionet.org/content/multimodal-emergency-benchmark/

Journal ref J.M. Lopez Alcaraz, H. Bouma, N. Strodthoff, Enhancing clinical decision support with physiological waveforms -- A multimodal benchmark in emergency care, Computers in Biology and Medicine, Vol. 192, Part A, 2025, 110196

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12219 2024-10-17 cs.AI cs.CL cs.MM 80%

OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Lichang Chen, Hexiang Hu, Mingda Zhang, Yiwen Chen, Zifeng Wang, Yandong Li, Pranav Shyam, Tianyi Zhou, Heng Huang, Ming-Hsuan Yang, Boqing Gong

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 19 pages, 6 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.10504 2021-09-06 eess.IV 80%

Liver Segmentation from Multimodal Images using HED-Mask R-CNN

Supriti Mulay, Deepika G, Jeevakala S, Keerthi Ram, Mohanasankar Sivaprakasam

专题命中 多模态评测 :multimodal(title,abstract)

Comments Accepted in 1st International Workshop on Multiscale Multimodal Medical Imaging (MMMI 2019) - MICCAI 2019

Journal ref Multiscale Multimodal Medical Imaging. MMMI 2019. Lecture Notes in Computer Science, vol 11977

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21497 2026-08-25 eess.IV cs.CV 新提交 79%

CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets

CHIMERA挑战赛:利用多模态数据集预测前列腺癌患者的生化复发

Robert N. Spaans, Catherine Chia, Tongjie Wang, Adam Kowalewski, Parandzem Khachatryan, Domingos Oliveira, Khrystyna Faryna, Jean-Paul A. van Basten, Geert Litjens, Nadieh Khalili

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究开发首个公开标准化前列腺癌预后多模态基准CHIMERA挑战赛,整合多类数据建立数据集,发现多模态模型在缺失临床变量时更稳健,单模态临床模型性能更高但对变量完整性敏感。

Comments 38 pages, 3 figures, 3 supplementary figures. Preprint submitted to Medical Image Analysis. Challenge results presented at the CHIMERA workshop, MICCAI 2025. Challenge website: https://chimera.grand-challenge.org/chimera/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22713 2026-08-25 cs.CL 新提交 79%

A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports

基于来源的从临床病例报告构建和评估渐进式多模态诊断对话的框架

Yufan Wang, Rui Yang, Yi Liu, Yi Lin, Yifan Peng

机构 * Weill Cornell Medicine(威尔康奈尔医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 该研究提出基于来源的框架用于构建渐进式多模态诊断对话,并评估MLLMs的诊断推理能力,实验显示所提框架转换病例报告的参考对话表现优异,而前沿MLLMs的相关指标显著更低。

Comments Accepted to IEEE HealthCom 2026, Distinguished Invited Papers Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22323 2026-08-25 cs.CV 新提交 79%

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

MedReaMM:评估大型多模态模型在专家级临床诊断综合能力上的表现

Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai

机构 * The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院) Jiangsu Province Hospital(江苏省医院) Huazhong University of Science Technology(华中科技大学) Fudan University(复旦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究构建了多模态临床诊断基准MedReaMM,评估23个大型多模态模型的诊断综合能力,发现多数模型准确率不足50%,揭示了该领域的能力差距及相关影响因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21883 2026-08-25 cs.CV 新提交 79%

VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

VIG:作为多模态思维链压缩奖励信号的视觉信息增益

Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究提出VIG奖励机制,通过提升视觉信息密度优化多模态CoT的准确率与效率权衡,无需额外资源,在多类基准及不同规模模型上均有效。

Comments Accepted by EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18425 2026-08-25 cs.CL 79%

Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs

多模态任务干扰:多模态大语言模型中历史-目标不匹配的基准与分析

Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 研究多模态大语言模型中任务切换导致的性能下降问题,通过六个任务的基准测试发现,从纯文本切换到图像目标会导致严重性能下降,而反向切换则影响较小,且多模态差异是主要干扰因素。

Journal ref Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 9282-9290, Palma de Mallorca, Spain. ELRA Language Resource Association

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20874 2026-08-24 cs.CV cs.RO 新提交 79%

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

面向自动驾驶的结合语义属性的多模态交通标志检测

Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani

机构 * Arriver System Software S.r.l.(Arriver系统软件有限责任公司) Qualcomm Auto Ltd Sweden Filial(高通汽车有限公司瑞典分公司) Qualcomm Auto Ltd.(高通汽车有限公司) Qualcomm Technologies, Inc(高通技术公司)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对自动驾驶中交通标志检测的跨区域泛化差、远距离小目标检测弱、跟踪易受透视畸变影响的问题,提出多模态框架,结合LiDAR与相机,引入强度感知可变形融合模块、双运动模型跟踪器及语义属性分类流水线,在多国数据集上取得低漏检率,实现全球可泛化的商用级交通标志感知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20384 2026-08-24 cs.AI 新提交 79%

Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles

基于线性判别树集成的可解释多模态分类

Mojtaba Moattari

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究针对多模态分类器需平衡准确率与可解释性的需求,提出基于线性判别树集成的框架,在多模态情感行为分类任务中实现了优于基线模型的性能与更高的人类标注一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20312 2026-08-21 cs.CV 新提交 79%

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Inter-X++:用于多模态人与人交互分析的综合基准

Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv, Yichao Yan, Xin Jin, Zhibo Chen, Xiaokang Yang, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(宁波工程学院) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究提出多模态人与人交互基准Inter-X++,构建含精细标注的大规模数据集,开发统一HHI框架OpenHHI,其在生成与感知任务上达最优性能,验证了统一表示的有效性。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19666 2026-08-21 cs.CV 新提交 79%

MUST-PET: MUltimodal Self-supervised learning across Tracers for whole-body PET/CT-based lesion segmentation

MUST-PET:用于基于全身PET/CT的病灶分割的跨示踪剂多模态自监督学习

Bashirul Azam Biswas, Amartya Bhattacharya, Biratal Raj Wagle, Matthew E. Maeder, James B. Yu, Indrani Bhattacharya

机构 * Geisel School of Medicine at Dartmouth(达特茅斯盖泽尔医学院) Dartmouth Hitchcock Medical Center(达特茅斯-希区柯克医疗中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出MUST-PET多模态自监督学习框架,通过跨示踪剂的上下文感知掩码重建,实现全身PET-CT病灶的标注高效、可泛化分割,性能优于从头训练模型。

Comments Submitted to SPIE CAD 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14497 2026-08-21 cs.CV 版本更新 79%

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

通过自我场景增强在多模态大语言模型中强化自我中心空间感知

Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究如何强化多模态大语言模型的自我中心空间感知,提出自我场景增强框架ESA,利用自我元素图作为中间表示,通过视觉基础模型增强空间感知,在EgoTextVQA基准上取得显著性能提升。

Comments 14 pages, 8 figures. Chi Kit Wong and Ye Pan contributed equally. Code: https://github.com/Chikit-WONG/spatialGraph

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18096 2026-08-20 cs.CL cs.LG 新提交 79%

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

MAVEN:基于紧凑对齐评估器的多模态内容宏观社会价值评估框架

Zijuan Zhao, Zheren Fu, Hou Xia, Licheng Zhang, Yi Liu, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 针对多模态内容宏观社会价值评估难题,提出MAVEN分层框架,构建相关基准与指标,优化评估器,实验表明其2B评估器表现接近前沿闭源VLMs,提供了可扩展评估路径。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15265 2026-08-20 cs.AI 版本更新 79%

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

VibeWorlding:多模态智能体能否端到端构建3D开放世界?

Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu

机构 * HKUST(GZ)(香港科技大学(广州)) Tencent(腾讯) AI Thrust TEG AIPD

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究提出VibeWorlding框架,构建VWE-BENCH基准与VibeWorlding-Gym框架,发现当前前沿MLLMs在3D开放世界构建任务中表现不佳,经RL训练的VibeWorlder模型可提升性能,旗舰模型VibeWorlder-30B-A3B表现最优。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17809 2026-08-20 cs.CV 版本更新 79%

Million-scale multimodal pollen microscopy with expert-guided foundation models

百万级多模态花粉显微镜图像与专家引导的基础模型

András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai

机构 * Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系) The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室) National Centre for Public Health and Pharmacy(国家公共卫生与药品中心) INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区) National Korányi Institute for Pulmonology(国家科拉尼肺病研究所) Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心) Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。

Comments 31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports. v2: clarified reporting of taxonomic scope, captioning settings, backbone configuration, and evaluation details; no changes to numerical results or conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13068 2026-08-20 cs.CV cs.LG 版本更新 79%

Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

聚焦图像:用于胸部X光诊断与报告生成的注视监督多模态学习

Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy, Farig Sadeque, Swakkhar Shatabda

机构 * Department of Computer Science and Engineering, Brac University(计算机科学与工程系,布拉克大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究基于MIMIC-Eye数据集构建两阶段多模态框架,引入注视监督提升胸部X光诊断准确率与报告空间可解释性,为该领域提供了可复现的新基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17884 2026-08-19 cs.CV 新提交 79%

CFB-GBM v2.0: An Augmented Longitudinal Dataset for Multi-Modal Glioblastoma Segmentation, Radiomics, and RANO Progression Tracking

CFB-GBM v2.0:用于多模态胶质母细胞瘤分割、放射组学及RANO进展追踪的增强纵向数据集

Alexandre G. Leclercq, Noémie N. Moreau, Hugo Audebert, Andros Nassar, Thomas Cochin, Thomas Leleu, Loïc Le Henaff, Alexis Desmonts, Yoann Poirier, Aurélie Dubru, Laura Guillemette, Pascal Lecoeur, Kévin Lemasson, Cyril Jaudet, Sébastien Bougleux, Romain Hérault, Carole Brunaud, Samuel Valable, Dinu Stefan, Charlotte Raboutet, Alain Batalla, Joëlle Lacroix, Roman Rouzier, Aurélien Corroyer-Dulmont

机构 * Centre François Baclesse(弗朗索瓦·巴克莱斯中心) Université de Caen Normandie(卡昂诺曼底大学) ENSICAEN(卡昂高等工程师学院) GREYC(格雷计算机科学研究中心)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文发布增强版CFB-GBM v2.0纵向数据集,含264名GBM患者数据,完成所有时间点GTV勾画(完成率达97%),提供相关标注、特征及WHO分类信息,可用于多模态GBM相关研究。

Comments 9 pages, 2 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08093 2026-08-19 cs.AI 版本更新 79%

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

面向证据基础计算病理学的多模态智能体协同助手

Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Yihui Wang, Jiabo Ma, Ling Liang, Yingxue Xu, Zhengrui Guo, Guanghao Wu, Danyi Li, Ziqi Zhou, Donglin Tan, Zhijian Cen, Ying Tan, Xiaolin Liu, Qi Xie, Xiaoying Tang, Xi Peng, Cheng Deng, Lijuan Qu, Ronald Cheong Kin Chan, Li Liang, Hao Chen

机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科) Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科) Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室) Jinfeng Laboratory(锦风实验室) Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出PathPocket,一种多模态AI协同助手,通过构建包含11万文档的病理证据语料库和455万实体的超图,实现基于证据的病理诊断,在20万真实案例上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00547 2026-08-19 cs.AI cs.LG 版本更新 79%

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

统一是否带来代价?Uni-SafeBench:一种用于统一多模态大模型的安全基准

Zixiang Peng, Yongxiu Xu, Qin-Yi Zhang, Jiexun Shen, Yi-Fan Zhang, Hongbo Xu, Yubin Wang, Gaopeng Gou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出Uni-SafeBench,用于评估统一多模态大模型的安全性,发现统一过程虽提升能力但显著降低基础LLM的安全性,开源资源以促进更安全的AGI发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16276 2026-08-18 cs.HC cs.CL 新提交 79%

PolyDebate: A Game-Orchestrated Multimodal System for Debate Skills Practice and Evaluation

PolyDebate:一种用于辩论技能练习与评估的游戏编排多模态系统

Jianing Yin, Weng Pan Kuan, Xiaoyun Liu, Zhiyuan Wen, Yuxuan Li, Milos Stojmenovic, Jiannong Cao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 PolyDebate是一款游戏化多模态辩论练习评估系统,含Unity 3D游戏与网页版本,经四项研究验证,可结合AI对手、多模态评估与结构化反馈助力学习者提升辩论技能。

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15999 2026-08-18 cs.AI 新提交 79%

MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment

MUPA²E:用于情绪评估的非对称注意力多模态统一感知框架

Stefanos Gkikas, Eric Nichols, Christian Arzate Cruz, Randy Gomez

机构 * Honda Research Institute Japan(日本本田研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出MUPA²E多模态统一感知框架,用共享非对称注意力骨干处理面部视频与EEG,在DMER数据集上对比多配置,发现时长线索影响分类,控制时长后准确率下降,证明统一架构处理异质信号的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12407 2026-08-18 cs.RO cs.CV cs.LG 版本更新 79%

MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery

MiDAS:一种用于机器人辅助微创手术的多模态数据采集系统和数据集

Keshara Weerasinghe, Seyed Hamid Reza Roodabeh, Andrew Hawkins, Zhaomeng Zhang, Zachary Schrader, Homa Alemzadeh

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 MiDAS是一种开源多模态数据采集系统,能够实现机器人辅助微创手术的非侵入式数据采集,并发布首个高保真仿真模型的缝合任务数据集。

Comments 29 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05923 2026-08-17 eess.SP cs.AI cs.LG eess.IV q-bio.QM 79%

Cedalion Tutorial: A Python-based framework for comprehensive analysis of multimodal fNIRS & DOT from the lab to the everyday world

Cedalion教程:一个基于Python的框架,用于对多模态fNIRS与DOT进行综合分析,从实验室到日常世界

E. Middell, L. Carlton, S. Moradi, T. Codina, T. Fischer, J. Cutler, S. Kelley, J. Behrendt, T. Dissanayake, N. Harmening, M. A. Yücel, D. A. Boas, A. von Lühmann

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 Cedalion是一个基于Python的开源框架,用于统一多模态fNIRS和DOT数据的分析,支持可重复、可扩展的神经成像工作流程。

Comments 33 pages main manuscript, 180 pages Supplementary Tutorial Notebooks, 12 figures, 6 tables, under review in SPIE Neurophotonics

Journal ref Neurophotonics, Vol. 13, Issue S3, S32602 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏