arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9251 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9251 篇

2604.22835 2026-04-28 cs.CV cs.AI 62%

ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes

ParkingScenes: 一个用于模拟场景中端到端自动驾驶泊车的结构化数据集

Haonan Chen, Kaiwen Xiao, Bin Tian, Jun Fu

机构 * Institute of Automation(自动化研究所) Chinese Academy of Sciences(中国科学院) School of Art and Design(艺术与设计学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ParkingScenes数据集,通过结构化轨迹生成提升泊车性能,展示结构化监督对自主泊车政策学习的有效性。

Comments Accepted by CAC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22822 2026-04-28 cs.CV cs.AI 62%

DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models

DO-Bench: 一种用于诊断视觉语言模型中物体幻觉的可归因基准

JiYang Wang, Jiawei Chen, Mengqi Xiao, Yu Cheng, Yangfu Li, Zhaoxia Yin

机构 * Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University(多维信息处理上海市重点实验室,东华大学) Zhongguancun Academy(中关村学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DO-Bench,通过结构化多模态干预隔离错误来源,评估模型对文本先验和感知能力的鲁棒性,揭示物体幻觉的异质性失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22367 2026-04-27 cs.CL cs.AI 62%

CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

CNSL-bench:用于评估多模态大语言模型在中文国家手语理解能力的基准

Rui Zhao, Xuewen Zhong, Xiaoyun Zheng, Jinsong Su, Yidong Chen

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China(福建省-台湾非物质文化遗产数字化保护与智能处理重点实验室(XMU),文化和旅游部,中国) National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China(教育与教学媒体语言资源监测与研究中心,厦门大学,中国)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CNSL-bench,首个评估多模态大语言模型在中文国家手语理解能力的基准,通过权威 grounding、多模态覆盖和手部动作多样性,评估21个模型,发现当前模型在手语理解上仍显著劣于人类表现。

Comments Accepted as the Main Conference at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22333 2026-04-27 cs.CV cs.AI 62%

ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding

ChangeQuery: 从视觉检测到语义理解,推动自然灾害和人为灾害的遥感变化分析

Dongwei Sun, Jing Yao, Kan Wei, Xiangyong Cao, Chen Wu, Zhenghui Zhao, Pedram Ghamisi, Jun Zhou, Jón Atli Benediktsson

机构 * School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学) State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences(遥感与数字地球国家重点实验室,航空信息研究所,中国科学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克研究所) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Faculty of Electrical and Computer Engineering, University of Iceland(冰岛大学电气与计算机工程学院) School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ChangeQuery通过多模态框架实现灾害情境感知,结合DICQ数据集和自动化标注管道,提升灾害监测的交互性和解释性,实现高精度的灾害量化与总结。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22260 2026-04-27 cs.CV cs.AI 62%

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

迈向安全的移动性:由开放-ended视觉-语言数据集驱动的统一运输基础模型

Wenhui Huang, Songyan Zhang, Collister Chua, Yang Liang, Zhiqi Mao, Heng Yang, Chen Lv

机构 * Nanyang Technological University(南洋理工大学) Harvard University(哈佛大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Land Transportation Dataset数据集和UniVLT模型,通过开放-ended视觉问答和多任务学习,提升城市交通环境下的安全推理能力,实现微观自动驾驶与宏观交通分析的统一。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21144 2026-04-24 cs.CL cs.AI cs.HC 62%

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

利用机器心智 imagery 代表情境对话中的共同背景

Biswesh Mohapatra, Giovanni Duca, Laurent Romary, Justine Cassell

机构 * Inria University of Trento(特伦托大学) Inria, Carnegie Mellon University(Inria 和卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种主动视觉支架框架,通过将对话状态转化为持久的视觉历史,减少代表模糊并提高对话理解。

Comments Work under review. Biswesh Mohapatra and Giovanni Duca both contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08620 2026-04-22 cs.AI cs.CV 62%

ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios

ViDoRe V3:对复杂现实场景中检索增强生成的综合评估

António Loison, Quentin Macé, Antoine Edy, Victor Xing, Tom Balough, Gabriel Moreira, Bo Liu, Manuel Faysse, Céline Hudelot, Gautier Viaud

机构 * Illuin Technology(Illuin技术公司) NVIDIA CentraleSupélec, Paris-Saclay(巴黎萨克雷中央理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ViDoRe V3是一个多模态检索增强生成基准,涵盖10个专业领域数据集,通过12000小时人工标注评估先进RAG管道,发现视觉检索优于文本检索,晚期交互模型和文本重排序显著提升性能,但模型在非文本元素、开放性查询和细粒度视觉定位上仍有不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16822 2026-04-22 cs.CV cs.AI 62%

ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition

ReefNet:一个大规模数据集和基准,用于细粒度珊瑚礁识别

Abdulwahab Felemban, Yahia Battach, Faizan Farooq Khan, Yuqian Fu, Xuhui Liu, Yesmeen M. Khattab, Yousef A. Radwan, Xiang Li, Fabio Marchese, Sara Beery, Burton H. Jones, Francesca Benzoni, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology (KAUST)(国王阿卜杜勒·阿齐兹科技大学) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ReefNet,一个大规模珊瑚礁图像数据集,用于细粒度识别。通过专家验证和过滤,构建了高置信度基准子集,揭示了多模态模型在生物多样性监测中的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07966 2026-04-22 cs.CV cs.CL 62%

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

视觉-表格问答:面向表格图像推理的开放领域基准

Boammani Aser Lompo, Marc Haraoui

机构 * École de Technologie Supérieure(埃克塞技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出Visual-TableQA,一个大规模开放领域多模态数据集,用于评估和提升视觉推理能力,包含2500个LaTeX渲染表格和6000对推理密集型问答对,通过多模型协作生成,验证了模型在外部基准上的鲁棒性。

Comments Accepted at the First Workshop on Foundations of Reasoning in Language Models, NeurIPS 2025. Available at: https://openreview.net/forum?id=fvJRsGwhPf

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03331 2026-04-21 cs.CV cs.AI cs.LG 62%

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

MMErroR:面向视觉-语言模型错误推理的基准测试

Yang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu, Mingxuan Huang, Jingchao Wang, Zhihong Zhu, Boyan Xu, Zhiqi Huang

机构 * Guangdong University of Technology(广东工业大学) Hong Kong Baptist University(香港 Baptist大学) Sun Yat-sen University(中山大学) Peking University(北京大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MMErroR基准,通过1997个包含单一推理错误的样本,评估视觉-语言模型检测错误推理的能力,发现即使最佳模型也仅能正确分类66.65%的错误。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17570 2026-04-21 cs.CV cs.AI 62%

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

PBSBench:一种多级视觉-语言框架和血液病理学全滑片图像解释基准

Yuanlong Wang, Weichi Chen, Adrian Rajab, Wenfang Liu, Yulan Jin, Andrew Srisuwananukorn, Ping Zhang

机构 * The Ohio State University(俄亥俄州立大学) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PBSBench,一种针对血液病理学全滑片图像解释的多级视觉-语言框架和基准,通过构建PBSInstr数据集和PBS-VL模型,提升了对血涂片图像的理解能力。

Comments 19 pages, 12 figures, Accepted by CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16538 2026-04-21 cs.CV cs.CL 62%

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

VC-Inspector:通过事实分析推进无参考视频字幕评估

Shubhashis Roy Dipta, Tz-Ying Wu, Subarna Tripathi

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校) Intel(英特尔)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 VC-Inspector提出了一种轻量级开源大多模态模型,用于无参考视频字幕评估,专注于事实准确性。通过生成可控制事实错误的字幕及评分注释,实现了可重复和事实感知的评估,实验显示其与人类判断的高相关性。

Comments Accepted at ACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16851 2026-04-21 cs.LG cs.AI cs.CV q-bio.BM q-bio.QM 62%

Applications of deep generative models to DNA reaction kinetics and to cryogenic electron microscopy

深度生成模型在DNA反应动力学和低温电子显微镜中的应用

Chenwei Zhang

机构 * The University Of British Columbia(不列颠哥伦比亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文探讨了深度生成模型在DNA反应动力学和低温电子显微镜分析中的应用,提出ViDa和Struc2mapGAN等方法,提升生物过程的可解释性和结构建模的准确性。

Comments PhD Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16506 2026-04-21 cs.CV cs.CL 62%

Medical thinking with multiple images

多图像医学推理

Zonghai Yao, Benlu Wang, Yifan Zhang, Junda Wang, Iris Xia, Zhipeng Tang, Shuo Han, Feiyun Ouyang, Zhichao Yang, Arman Cohan, Hong Yu

机构 * Manning College of Information and Computer Sciences, UMass Amherst(UMass阿默斯特信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗保健中心) Department of Computer Science, Yale University(耶鲁大学计算机科学系) Miner School of Computer and Information Sciences, UMass Lowell(UMass洛威计算机与信息科学学院) Department of Electrical and Computer Engineering, UMass Lowell(UMass洛威电气与计算机工程系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出MedThinkVQA多图像医学推理基准,通过专家标注数据验证多图像整合对临床推理的重要性,发现多图像推理瓶颈在于证据提取与对齐,且增加推理时间计算效果有限。

Comments Equal contribution for the first two authors. To appear in the proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026). Code is in https://github.com/benluwang/MedThinkVQA. Dataset is in https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16104 2026-04-20 eess.IV cs.AI cs.CV 62%

Dual-Modal Lung Cancer AI: Interpretable Radiology and Microscopy with Clinical Risk Integration

双模肺癌AI:结合可解释放射学与显微镜的临床风险整合

Baramee Sukumal, Aueaphum Aueawatthanaphisut

机构 * Hatyaiwittayalai School(哈雅维泰莱学校) School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院) Sirindhorn International Institute of Technology(Sirindhorn国际技术学院) Thammasat University(泰国朱拉大学) Pathum Thani, Thailand(帕亨他尼,泰国)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出整合CT放射学与H&E病理学的双模AI框架,利用卷积神经网络提取特征并融合临床数据,通过可解释AI技术提升诊断性能,实现肺癌亚型分类。

Comments 16 pages, 6 figures, 3 tables, 8 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09836 2026-04-20 cs.AI cs.CL cs.LG 62%

COMPOSITE-Stem

COMPOSITE-STEM基准:通过70个专家编写的任务评估AI代理的科学推理能力

Kyle Waters, Lucas Nuzzi, Tadhg Looram, Alessandro Tomasiello, Ariel Ghislain Kemogne Kamdoum, Bikun Li, Damien Sileo, Egor Kretov, Francesco Fournier-Facio, Georgios Soloupis, Haile Kassahun, Hew Wolff, Jiaqi Cai, Lianghui Li, Marc Roth, Mohinder Naiya, Naixu Guo, Qicheng Tang, Richard Wheeler, Samuele Sala, Serguei Popov, Steven Dillmann, Yuqi Li

机构 * PortexAI University of Milano-Bicocca(米兰-比科卡大学) University of Calgary(卡尔加里大学) University of Chicago(芝加哥大学) Inria(法国国家信息与自动化技术研究院) Fraunhofer Institute for Individualized Medical Technology IMTE(弗劳恩霍夫个性化医疗技术研究所) University of Cambridge(剑桥大学) Independent(独立) McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) Queen Mary University of London(伦敦大学玛丽女王学院) Dot Ingredients National University of Singapore(新加坡国立大学) Georgia Institute of Technology(佐治亚理工学院) University of Edinburgh(爱丁堡大学) Murdoch University(墨尔本大学) University of Porto(波尔图大学) Stanford University(斯坦福大学) Stony Brook University(史泰津布鲁克大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出COMPOSITE-STEM基准,通过70个专家编写任务评估AI代理的科学推理能力,结合精确匹配评分和LLM作为评委的协议,展示AI在科学领域的能力突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21675 2026-04-20 cs.CL cs.CV cs.GR 62%

Is this chart lying to me? Automating the detection of misleading visualizations

这张图表在欺骗我吗?自动化检测误导性可视化

Jonathan Tonglet, Jan Zimny, Tinne Tuytelaars, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系、图腾斯大学(TU Darmstadt)及应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(电子工程系、鲁文大学) Department of Computer Science, KU Leuven(计算机科学系、鲁文大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 研究提出Misviz基准测试集和合成数据集,用于自动化检测误导性可视化,发现该任务仍极具挑战性。

Comments Camera-ready version accepted at ACL 2026 Main conference. Code and data available at: https://github.com/UKPLab/acl2026-misviz

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15066 2026-04-17 cs.LG cs.AI cs.MM 62%

Time-RA: Towards Time Series Reasoning for Anomaly Diagnosis with LLM Feedback

Time-RA:面向时间序列推理的异常诊断方法:基于LLM反馈

Yiyuan Yang, Zichuan Liu, Lei Song, Kai Ying, Zhiguang Wang, Tom Bamford, Svitlana Vyetrenko, Jiang Bian, Qingsong Wen

机构 * University of Oxford(牛津大学) Nanjing University(南京大学) MSRA(微软研究院) SJTU(上海交通大学) Abel AI Outsampler University of Strasbourg(斯特拉斯堡大学) Squirrel Ai Learning(Squirrel AI学习)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 本文提出Time-RA,通过引入RATs40K多模态数据集,改进时间序列异常检测的生成式推理方法,提升诊断准确性和解释性,实现可解释的多模态时间序列分析。

Comments ACL 2026 Findings. 27 pages, 11 figures, 15 tables. Code and dataset are publicly available

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10966 2026-04-17 cs.CV cs.AI 62%

You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass

你只评判一次:单次前向传递中的多响应奖励建模

Yinuo Yang, Zixian Ma, Manasi Ganti, Jieyu Zhang, Ranjay Krishna

机构 * Paul G. Allen School of Computer Science & Engineering(保罗·G·艾伦计算机科学与工程学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种判别性多模态奖励模型,通过单次前向传递对所有候选响应进行评分。该方法通过拼接多个响应并应用交叉熵实现直接比较推理和高效N-way偏好学习,同时提升速度和计算效率。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13942 2026-04-16 cs.CV cs.AI cs.LG 62%

Frozen Forecasting: A Unified Evaluation

冻结预测:一种统一的评估

Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera, Guangyao Zhou, Sayna Ebrahimi, Rishabh Kabra, Carl Doersch, Maks Ovsjanikov, João Carreira, Shiry Ginosar

机构 * Google DeepMind(谷歌DeepMind) Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种统一的评估框架,用于评估冻结视觉骨干在多样化任务和抽象层次上的预测能力,通过训练潜在扩散模型直接在表示空间中预测未来特征,发现预测性能与感知质量密切相关,视频预训练模型表现优于图像预训练模型。

Comments New Title, Additional Author

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19088 2026-04-16 cs.CL cs.CV 62%

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

破解 juxtaposition 的密码:AI 模型能否理解幽默的矛盾

Zhe Hu, Tuo Liang, Jing Li, Yiren Lu, Yunlai Zhou, Yiran Qiao, Jing Ma, Yu Yin

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文探讨了AI在理解基于矛盾叙事的幽默中的挑战,通过引入YesBut基准测试,评估大型语言模型在识别和解释漫画中的表现,发现即使是最先进的模型仍无法匹敌人类水平。

Comments NeurIPS 2024 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13029 2026-04-15 cs.CV cs.AI 62%

Visual Preference Optimization with Rubric Rewards

基于评分标准的视觉偏好优化

Ya-Qi Yu, Fangyu Hong, Xiangyang Qu, Hao Wang, Gaojie Wu, Qiaoyu Luo, Nuo Xu, Huixin Wang, Wuheng Xu, Yongxin Liao, Zihao Chen, Haonan Li, Ziming Li, Dezhi Peng, Minghui Liao, Jihao Wu, Haoyu Ren, Dandan Tu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出rDPO框架,通过实例特定的评分标准提升多模态任务中的偏好优化效果,实验表明其在多个基准测试中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12512 2026-04-15 cs.CV cs.AI 62%

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)

NTIRE 2026 第三次恢复任何图像模型(RAIM)挑战:专业图像质量评估(Track 1)

Guanyi Qin, Jie Liang, Bingbing Zhang, Lishen Qu, Ya-nan Guan, Hui Zeng, Lei Zhang, Radu Timofte, Jianhui Sun, Xinli Yue, Tao Shao, Huan Hou, Wenjie Liao, Shuhao Han, Jieyu Yuan, Chunle Guo, Chongyi Li, Zewen Chen, Yunze Liu, Jian Guo, Juan Wang, Yun Zeng, Bing Li, Weiming Hu, Hesong Li, Dehua Liu, Xinjie Zhang, Qiang Li, Li Yan, Wei Dong, Qingsen Yan, Xingcan Li, Shenglong Zhou, Manjiang Yin, Yinxiang Zhang, Hongbo Wang, Jikai Xu, Zhaohui Fan, Dandan Zhu, Wei Sun, Weixia Zhang, Kun Zhu, Nana Zhang, Kaiwei Zhang, Qianqian Zhang, Zhihan Zhang, William Gordon, Linwei Wu, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Cici Liu, Yaokun Shi

机构 * NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)(NTIRE 2026 第三次恢复任何图像模型(RAIM)挑战:专业图像质量评估(Track 1))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出NTIRE 2026挑战赛,聚焦专业图像质量评估,通过多模态大语言模型提升图像评估的准确性与解释性,吸引200多团队参与,推动专业IQA领域的发展。

Comments NTIRE Challenge Report. Accepted by CVPRW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17012 2026-04-14 cs.CV cs.AI 62%

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

SpatialScore:迈向空间智能的综合评估

Haoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SpatialScore,首个全面评估多模态空间智能的基准,评估49个模型发现其在空间理解上存在显著差距,并构建了SpatialCorpus和SpatialAgent提升模型性能。

Comments Accepted by CVPR 2026 (Highlight); Project Page: https://haoningwu3639.github.io/SpatialScore

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11004 2026-04-14 cs.CV cs.AI cs.LG 62%

Panoptic Pairwise Distortion Graph

全景成对失真图

Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi

机构 * Huawei Technologies(华为技术有限公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出全景成对失真图任务,通过区域结构化组成图像对,改进现有方法对整体图像分析的局限性,构建PandaSet数据集、PandaBench基准和Panda架构,揭示多模态大语言模型在区域级失真理解上的不足。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10666 2026-04-14 cs.CV cs.CL cs.LG 62%

Omnimodal Dataset Distillation via High-order Proxy Alignment

通过高阶代理对齐实现多模态数据集蒸馏

Yuxuan Gao, Xiaohao Liu, Xiaobo Xia, Tongliang Liu

机构 * University of Science and Technology of China(中国科学技术大学) Xidian University(西安电子科技大学) National University of Singapore(新加坡国立大学) The University of Sydney(悉尼大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出HoPA方法,通过紧凑代理捕捉高阶跨模态对齐,解决多模态数据集蒸馏中的异质性和复杂跨模态交互问题,实现可扩展的联合蒸馏。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10485 2026-04-14 cs.CV cs.AI 62%

UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

UDAPose:无监督领域适应用于低光人类姿态估计

Haopeng Chen, Yihao Ai, Kabeen Kim, Robby T. Tan, Yixin Chen, Bo Wang

机构 * University of Mississippi(密西西比大学) National University of Singapore(新加坡国立大学) Duksung Women’s University(德成女子大学) ASUS Intelligent Cloud Services (AICS)(华硕智能云服务(AICS))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 UDAPose提出了一种新的框架,通过合成低光图像并动态融合视觉线索与姿态先验,提升低光环境下的人体姿态估计性能。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09876 2026-04-14 cs.LG cs.AI cs.CV cs.HC 62%

Efficient Personalization of Generative User Interfaces

高效生成用户界面的个性化

Yi-Hao Peng, Samarth Das, Jeffrey P. Bigham, Jason Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了生成用户界面个性化难题,提出基于历史设计师偏好样本的高效方法,在技术评估中优于预训练模型,并能生成更多用户偏好界面。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08615 2026-04-13 cs.CV cs.AI 62%

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments

MARINER:一种基于3E驱动的基准,用于开放水域环境中的细粒度感知与复杂推理

Xingming Liao, Ning Chen, Muying Shu, Yunpeng Yin, Peijian Zeng, Zhuowei Wang, Nankai Lin, Lianglun Cheng

机构 * Guangdong University of Technology(广东工业大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MARINER基准,通过3E框架评估开放水域环境中的细粒度感知与复杂推理,揭示先进模型在复杂海洋场景中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07883 2026-04-10 cs.AI cs.CL cs.CY cs.MA 62%

An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks

用于教育教科书历史偏见检测的代理评估架构

Gabriel Stefan, Adrian-Marius Dumitran

机构 * University of Bucharest(布加勒斯特大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种代理评估架构,通过多模态筛查代理、异质陪审团和元代理,有效检测教育教科书中的隐性偏见,实验表明其在经济性和准确性上具有优势。

Comments Accepted for ITS(Intelligent Tutoring Systems) 2026 Full Paper

详情

展开后加载摘要…

URL PDF HTML 收藏