arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2605.02206 2026-05-11 cs.CV cs.LG 79%

Metric Unreliability in Multimodal Machine Unlearning: A Systematic Analysis and Principled Unified Score

多模态机器去学习中的度量不可靠性:系统分析与原则统一分数

Abdullah Ahmad Khan, Hamid Laga, Ferdous Sohel

机构 * Murdoch University(墨尔本大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文系统分析了多模态去学习中度量可靠性问题,提出统一质量评分(UQS)以解决不同度量标准间的冲突,通过实验验证了UQS在稳定排序方面的有效性。

Comments 9 Pages , 6 figures, Neurips 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01333 2026-05-11 cs.CL 79%

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

OralMLLM-Bench: 评估多模态大语言模型在牙科实践中的认知能力

Rongyang Wang, Shuang Zhou, Jiashuo Wang, Wenya Xie, Xiaoxia Che

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出OralMLLM-Bench,通过27个临床任务评估多模态大语言模型在牙科影像分析中的认知能力,揭示模型与牙科医生的性能差异及改进方向。

Comments 21 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15884 2026-05-11 cs.CV 79%

Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching

Contrast-X:一个多模态对比图像合成基准及通用模态流匹配

Yifan Chen, Fei Yin, Hao Chen, Jia Wu, Chao Li

机构 * University of Cambridge(剑桥大学) MD Anderson Cancer Center(MD安德森癌症中心) University of Dundee(邓迪大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 Contrast-X通过多器官CT和乳腺DCE-MRI数据,提出FlowMI模型解决多模态缺失问题,评估合成图像质量及跨器官泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06584 2026-05-08 cs.AI 79%

NeuroAgent: LLM Agents for Multimodal Neuroimaging Analysis and Research

NeuroAgent:用于多模态神经影像分析和研究的LLM代理

Lujia Zhong, Yihao Xia, Jianwei Zhang, Shuo huang, Jiaxin Yue, Mingyang Xia, Yonggang Shi

机构 * Stevens Neuroimaging and Informatics Institute, Keck School of Medicine, University of Southern California(史蒂文斯神经影像与信息学研究所,凯克医学院,南加州大学) Viterbi School of Engineering, University of Southern California(维特比工程学院,南加州大学) Alfred E. Mann Department of Biomedical Engineering, Viterbi School of Engineering, University of Southern California(阿尔弗雷德·E·曼生物医学工程部门,维特比工程学院,南加州大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 NeuroAgent通过LLM驱动的代理框架自动化多模态神经影像预处理与分析,支持自然语言查询,提升神经影像研究的自动化与可重复性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06371 2026-05-08 cs.AI 79%

Debiased Multimodal Personality Understanding through Dual Causal Intervention

通过双因果干预实现去偏的多模态人格理解

Yangfu Zhu, Zitong Han, Nianwen Ning, Yuting Wei, Yuandong Wang, Hang Feng, Zhenzhou Shao

机构 * Capital Normal University(首都师范大学) Henan University(河南大学) University of International Relations(国际关系大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出双因果调整网络DCAN,通过因果模型消除多模态特征与人格特质间的偏见关联,提升预测准确性和公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06289 2026-05-08 stat.ML cs.AI cs.LG 79%

Multimodal Deep Generative Model for Semi-Supervised Learning under Class Imbalance

多模态深度生成模型用于类别不平衡下的半监督学习

Heegeon Yoon, Heeyoung Kim

机构 * Department of Industrial and Systems Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea(工业与系统工程系,韩国科学技术院(KAIST),大田,大韩民国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态深度生成模型,通过共享潜在变量和t分布改进类别不平衡问题,提升半监督学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05014 2026-05-08 cs.CV 79%

CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

CARD: 一种用于复杂道路地形密集3D重建的多模态汽车数据集

Gasser Elazab, Frank Neuhaus, Tilman Koß, Malte Splietker, Aditya Date, Michael Unterreiner, Maximilian Jansen, Olaf Hellwich

机构 * CARIAD SE(CARIAD公司) Technische Universität Berlin(柏林技术大学) Vision & Robotics GmbH(视觉与机器人技术有限公司)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 CARD数据集通过提供连续序列中的近密3D地面真实信息,解决了现有数据集在复杂道路地形下深度估计和补全的不足,支持更精确的几何和感知任务评估。

Comments Accepted at CVPR 2026 (Highlight). Project page: https://card.content.cariad.digital

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22859 2026-05-08 cs.CV 79%

From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models

从盲点到收益:面向大型多模态模型的诊断驱动迭代训练

Hongrui Jia, Chaoya Jiang, Yongrui Heng, Shikun Zhang, Wei Ye

机构 * Peking University(北京大学) Shandong University(山东大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DPE方法,通过诊断驱动的数据生成和强化学习,实现大型多模态模型的持续改进,实验表明在多个基准测试中取得稳定收益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05712 2026-05-08 cs.CV 79%

EgoEMG: A Multimodal Egocentric Dataset with Bilateral EMG and Vision for Hand Pose Estimation

EgoEMG:一个用于手姿态估计的多模态自体视角数据集,包含双侧EMG和视觉信息

Ziheng Xi, Jiayi Yu, Yitao Wang, Yanbo Duan, Jianjiang Feng, Jie Zhou

机构 * Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 EgoEMG数据集融合双侧EMG和视觉信息,用于手姿态估计,包含16通道EMG、IMU、RGB视频和RGB-D视频,涵盖41名参与者60种手势,提供EMG-to-pose、vision-to-pose和融合任务的基准测试。

Comments 34 pages, 13 figures, 15 tables. Submitted to NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05573 2026-05-08 astro-ph.IM cs.AI 79%

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification

AstroAlertBench: 评估多模态大语言模型在天文学分类中的准确性、推理和诚实性

Claire Chen, Jiabao Sean Xiao, Shuze Daniel Liu, Facundo Perez Paolino, Luke Handley, Theophile Jegou du Laz, Ricky Nilsson, Alice Zou, Matthew Graham, Ashish Mahabal

机构 * California Institute of Technology(加州理工学院) Massachusetts Institute of Technology(麻省理工学院) Purdue University(普渡大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出AstroAlertBench,通过多阶段逻辑链评估多模态大语言模型在天文学事件分类中的性能,揭示高精度与模型诚实性之间的矛盾,并引入人机协同评估协议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21471 2026-05-08 cs.AI 79%

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

SpatialBench: 多模态大语言模型空间认知的基准测试

Peiran Xu, Sudong Wang, Yao Zhu, Jianing Li, Gege Qi, Yunjian Zhang

机构 * Sun Yat-Sen University(中山大学) HKUST (GZ)(香港科技大学) Zhejiang University(浙江大学) Peking University(北京大学) CAICT(中国科学院电子技术研究所) UCAS(中国科学技术大学) CUC(中国科学技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出SpatialBench基准,通过五级空间认知框架评估多模态大语言模型的空间能力,揭示模型在感知与符号推理间的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01719 2026-05-08 cs.CL 79%

What MLLMs Learn about When they Learn about Multimodal Reasoning

大语言模型在学习多模态推理时所学的内容

Jiwan Chung, Neel Joshi, Pratyusha Sharma, Youngjae Yu, Vibhav Vineet

机构 * Microsoft Research AI Frontiers(微软研究院人工智能前沿) Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出MathLens基准测试,通过分解性能为感知、推理和多模态特定组件,揭示多模态推理模型评估中隐含的假设。研究显示,不同训练策略导致的能力谱系不同,多模态特定错误占比上升,表明进展反映的是子技能平衡变化而非统一提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01424 2026-05-05 cs.LG cs.AI 79%

Quantifying Multimodal Capabilities: Formal Generalization Guarantees in Pairwise Metric Learning

量化多模态能力:成对度量学习中的形式泛化保证

Richeng Zhou, Xuelin Zhang, Liyuan Liu

机构 * College of Informatics, Huazhong Agricultural Univeristy, Wuhan, China

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文通过理论分析多模态度量学习模型的泛化性质,揭示模态选择与算法性能的关系,推导新的泛化误差界,证明细粒度模态特征能降低假设空间复杂度,提升多模态学习系统的收敛速度和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28011 2026-05-01 cs.CV 79%

Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation

Echo-α:用于超声解读的大型代理多模态推理模型

Jing Zhang, Wentao Jiang, Tao Huang, Zhiwei Wang, Jianxin Liu, Jian Chen, Ping Ye, Gang Wang, Zengmao Wang, Bo Du, Dacheng Tao

机构 * MiliLab(Mili实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 Echo-α结合了精确的病变定位和整体临床推理,通过九任务监督课程和强化学习提升超声解读的准确性和可解释性。

Comments 12 pages, 4 figures. Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14602 2026-05-01 cs.RO cs.AI cs.HC 79%

K2MUSE: A human lower-limb multimodal walking dataset spanning task and acquisition variability for rehabilitation robotics

K2MUSE:一种涵盖任务和采集变异性的下肢多模态行走数据集,用于康复机器人

Jiwei Li, Bi Zhang, Xiaowei Tan, Wanxin Chen, Zhaoyuan Liu, Juanjuan Zhang, Weiguang Huo, Jian Huang, Lianqing Liu, Xingang Zhao

机构 * State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) College of Artificial Intelligence, Tianjin Key Laboratory of Intelligent Robotics, Nankai University(人工智能学院,天津智能机器人重点实验室,南开大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 K2MUSE数据集通过多模态数据揭示下肢康复机器人与生物力学信息的关系,为康复机器人开发提供大规模步态样本和数据驱动方法支持。

Comments 34 pages, 30 figures,7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27259 2026-05-01 cs.CV cs.LG 79%

VTBench: A Multimodal Framework for Time-Series Classification with Chart-Based Representations

VTBench: 一种基于图表表示的多模态时间序列分类框架

Madhumitha Venkatesan, Xuyang Chen, Dongyu Liu

机构 * University of California, Davis(加州大学戴维斯分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VTBench,通过融合原始序列和图表可视化,系统评估了时间序列分类中图表类型、视觉编码和数据集的影响,展示了图表模型在特定领域的竞争力及多模态融合的优势。

Comments 8 pages main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27217 2026-05-01 cs.AI 79%

Toward Personalized Digital Twins for Cognitive Decline Assessment: A Multimodal, Uncertainty-Aware Framework

迈向认知下降评估的个性化数字孪生:一种多模态、不确定性感知的框架

Bulent Soykan, Gulsah Hancerliogullari Koksalmis, Hsin-Hsiung Huang, Laura J. Brattain

机构 * Department of Mechanical, Industrial \& Manufac. Eng. The University of Toledo Toledo, OH, USA Department of Industrial Eng. Mngt. Systems University of Central Florida Orlando, FL, USA Department of Statistics Data Science University of Central Florida Orlando, FL, USA Department of Internal Medicine University of Central Florida Orlando, FL, USA

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种多模态、不确定性感知的框架,用于从稀疏、噪声和不规则的纵向数据中建模患者特定的疾病轨迹,通过结合潜在状态空间模型、多模态融合和不确定性感知验证,实现对认知下降的个性化评估。

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03583 2026-04-30 cs.MM cs.IR 79%

OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset

OpenLifelogQA:一个开放式多模态生活日志问答数据集

Quang-Linh Tran, Hoang-Bao Le, Tuong-Nghiem Diep, Binh Nguyen, Gareth J. F. Jones, Cathal Gurrin

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.MM

AI总结 OpenLifelogQA数据集包含14187对问答对,用于支持真实场景下的鲁棒评估,相比现有资源更适用于实际应用,通过评估LLaVA-NeXT-Interleave 7B模型展示了其在生活日志问答中的性能。

Comments In the proceedings of the 14th International Symposium on Information and Communication Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25618 2026-04-29 cs.MM 79%

Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding

超越孤立话语:基于提示的交互用于依赖上下文的对话多模态理解

Zhaoyan Pan, Hengyang Zhou, Xiangdong Li, Yuning Wang, Ye Lou, Jiatong Pan, Ji Zhou, Wei Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出CUCI-Net,通过结合局部模态证据与全局上下文证据,明确抽象上下文与话语之间的依赖关系,将其作为提示融入最终多模态交互阶段,实现上下文条件下的预测。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25072 2026-04-29 cs.CV 79%

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

超越准确率:统一多模态模型跨任务一致性的基准测试

Weixing Wang, Liudvikas Zekas, Anton Hackl, Constantin Alexander Auga, Parisa Shahabinejad, Jona Otholt, Antonio Rueda-Toicen, Gerard de Melo

机构 * Hasso Plattner Institute / University of Potsdam(霍普夫纳研究所 / 波茨坦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出XT-C-Bench,用于评估统一多模态模型跨任务的语义一致性,发现高生成或理解性能不等于强跨任务对齐,揭示一致性由模态间学习目标耦合紧密度决定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23799 2026-04-28 cs.CV 79%

VitaminP: cross-modal learning enables whole-cell segmentation from routine histology

VitaminP: 跨模态学习实现常规组织学图像的全细胞分割

Yasin Shokrollahi, Karina B. Pinao Gonzales, Elizve N. Barrientos Toro, Paul Acosta, Patient Mosaic Team, Pingjun Chen, Yinyin Yuan, Xiaoxi Pan

机构 * Department of Translational Molecular Pathology, Division of Pathology and Laboratory Medicine, The University of Texas MD Anderson Cancer Center(转化分子病理学部门,病理与实验室医学分会,德克萨斯大学MD安德森癌症中心) Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center(肿瘤数据科学研究所,德克萨斯大学MD安德森癌症中心) The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

AI总结 VitaminP通过跨模态学习从常规H&E图像实现全细胞分割,利用配对的H&E-mIF数据迁移分子边界信息以克服H&E的细胞质对比度限制,构建了跨模态监督策略,并在14个公开数据集上训练,超越了四种最新方法并在未见数据集上泛化。

Comments 44 pages, 10 figures. Code and models available

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10047 2026-04-28 cs.CV 79%

MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples

MuSc-V2:基于未标记样本互评分的零样本多模态工业异常分类与分割

Xurui Li, Feng Xue, Yu Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) University of Trento(特伦特大学)

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MuSc-V2框架,通过改进3D表示、融合多尺度特征及互评分机制,实现零样本多模态工业异常检测,提升MVTec 3D-AD和Eyecandies数据集性能。

Comments TPAMI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06667 2026-04-28 cs.CV cs.LG eess.IV 79%

Flood-DamageSense: Multimodal Mamba with Multitask Learning for Building Flood Damage Assessment using SAR Remote Sensing Imagery

Flood-DamageSense: 多模态Mamba与多任务学习用于基于SAR遥感图像的建筑物洪水损害评估

Yu-Hsuan Ho, Ali Mostafavi

机构 * Urban Resilience.AI Lab(城市韧性人工智能实验室) Zachry Department of Civil and Environmental Engineering(扎克里土木与环境工程系) Texas A&M University(德克萨斯A&M大学) College Station, TX(学院站, 德克萨斯)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Flood-DamageSense模型,结合多模态Mamba架构与多任务学习,利用SAR遥感图像实现建筑物洪水损害评估,通过融合前后事件影像与洪水风险层,提升洪水损害识别精度,实验显示在Hurricane Harvey数据集上F1值提升显著。

Journal ref Comput. Aided Civ. Infrastruct. Eng. 40.26 (2025) 4401-4424

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05231 2026-04-27 cs.RO cs.AI 79%

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction

Kaiwu:一种用于机器人学习和人机交互的多模态操控数据集和框架

Shuo Jiang, Haonan Li, Ruochen Ren, Yanmin Zhou, Zhipeng Wang, Bin He

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出Kaiwu多模态数据集,解决复杂装配场景中缺失的真实同步多模态数据问题,通过20名受试者和30个交互对象,记录11,664个整合动作实例,支持机器人学习、精细操作、人类意图研究和人机协作。

Comments 8 pages, 5 figures, Submitted to IEEE Robotics and Automation Letters (RAL)

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 11, pp. 11482-11489, Nov. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21575 2026-04-24 cs.CV cs.GR 79%

OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction

OmniFit:通过尺度无关密集地标预测实现多模态3D身体拟合

Zeyu Cai, Yuliang Xiu, Renke Wang, Zhijing Shao, Xiaoben Li, Siyuan Yu, Chao Xu, Yang Liu, Baigui Sun, Jian Yang, Zhenyu Zhang

机构 * Nanjing University(南京大学) Westlake University(西湖大学) iROOTECH Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 OmniFit通过尺度无关的密集地标预测,实现对多模态输入(如完整扫描、部分深度观测和图像)的无缝处理,首次在日常和宽松服装场景中超越多视图优化基线,达到毫米级精度。

Comments Project Page: https://zcai0612.github.io/OmniFit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21508 2026-04-24 cs.AI q-bio.BM 79%

BioMiner: A Multi-modal System for Automated Mining of Protein-Ligand Bioactivity Data from Literature

BioMiner:一种用于从文献中自动挖掘蛋白质-配体生物活性数据的多模态系统

Jiaxian Yan, Jintao Zhu, Yuhang Yang, Qi Liu, Kai Zhang, Zaixi Zhang, Xukai Liu, Boyan Zhang, Kaiyuan Gao, Jinchuan Xiao, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室) Princeton University(普林斯顿大学) Huazhong University of Science and Technology(华中科技大学) Infinite Intelligence Pharma(无限智能制药)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 BioMiner通过多模态框架提取生物活性数据,解决手动整理与文献增长不匹配的问题,其核心方法结合直接推理和化学结构解析,实现生物活性三元组的高准确率提取。

Comments 20 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20263 2026-04-23 q-bio.QM cs.AI cs.LG 79%

AROMA: Augmented Reasoning Over a Multimodal Architecture for Virtual Cell Genetic Perturbation Modeling

AROMA:增强的多模态架构下虚拟细胞基因扰动建模推理

Zhenyu Wang, Geyan Ye, Wei Liu, Man Tat Alexander Ng

机构 * AI for Life Sciences Lab, Tencent(腾讯AI生命科学实验室) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 AROMA通过整合文本证据、图拓扑信息和蛋白质序列特征,提升虚拟细胞基因扰动预测的准确性和可解释性,构建了包含498k样本的 PerturbReason 数据集,实验表明其在多个细胞系中表现优异。

Comments Accepted to ACL 2026 as a Findings paper. Zhenyu Wang and Geyan Ye are equal contributors; Geyan Ye is the corresponding author and project lead

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13232 2026-04-23 cs.AI cs.SE 79%

PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading

PlotChain: 多模态大语言模型在工程图表阅读中的确定性检查点评估

Mayank Ravishankara

机构 * Independent Researcher, San Francisco, CA, USA(独立研究者,美国加州旧金山)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 PlotChain提出了一种确定性的基于生成器的评估基准,用于评估多模态大语言模型在工程图表阅读中的性能,通过检查点实现子技能诊断和故障定位,展示了不同模型在图表阅读任务中的表现差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19567 2026-04-22 cs.AI 79%

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic

基于LLM的多模态推理用于视觉语义算术

Chuou Xu, Liya Ji, Qifeng Chen

专题命中 多模态评测 :multi-modal(title);cross-modal(abstract);分类 cs.AI

AI总结 本文提出两种新型任务和IRPD数据集,通过SAri-RFT方法提升大视觉语言模型的跨模态关系推理能力,实现视觉语义算术的突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02384 2026-04-22 cs.CV 79%

SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis

SMART-Ship:一个综合的同步多模对齐遥感目标数据集和基准,用于靠泊船舶分析

Chen-Chen Fan, Peiyao Guo, Linping Zhang, Kehan Qi, Haolin Huang, Yong-Qiang Mao, Yuxi Suo, Zhizhuo Jiang, Yu Liu, You He

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出SMART-Ship数据集,包含多模遥感图像,用于靠泊船舶分析,通过精细标注支持多模遥感任务,并评估了代表性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏