arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9205 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9205 篇

2509.05215 2026-04-13 cs.CL cs.LG 70%

BEDTime: A Unified Benchmark for Automatically Describing Time Series

BEDTime:一个统一的基准测试用于自动描述时间序列

Medhasweta Sen, Zachary Gottesman, Jiaxing Qiu, C. Bayan Bruss, Nam Nguyen, Tom Hartvigsen

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出BEDTime基准测试,评估模型对时间序列结构属性的描述能力,发现专为时间序列语言设计的模型表现欠佳,而视觉语言模型表现优异,且所有方法在现实鲁棒性测试中均表现脆弱。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00677 2026-04-02 cs.CV 70%

CL-VISTA: Benchmarking Continual Learning in Video Large Language Models

CL-VISTA:视频大语言模型持续学习基准测试

Haiyang Guo, Yichen Shi, Fei Zhu, Wenzhuo Liu, Hongbo Zhao, Fanhu Zeng, Shijie Ma, Da-Han Wang, Xu-Yao Zhang

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学前沿交叉科学学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FKLPRIU, School of Computer and Information Engineering, Xiamen University of Technology(厦门理工学院计算机与信息工程学院福建省模式识别与图像理解重点实验室)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 CL-VISTA通过8个多样化任务诱导分布偏移,评估持续学习方法在性能、效率和内存方面的权衡,揭示无单一最优解的特性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27183 2026-04-02 cs.CV 70%

Communicating about Space: Language-Mediated Spatial Integration Across Partial Views

关于空间的交流:通过语言介导的空间整合跨部分视图

Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil, Ponnurangam Kumaraguru, Aishwarya Agrawal

机构 * Mila – Quebec AI Institute(Mila – 魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) IIIT Hyderabad(海得拉巴国际信息技术学院)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 研究探讨多模态大语言模型能否通过对话整合不同视角,发现其在识别共享锚点物体上表现较好,但在关系推理和全局地图构建上存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16908 2026-04-02 cs.CV 70%

Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content

Q-REAL:迈向AI生成内容真实性和可信度评估

Shushi Wang, Zicheng Zhang, Chunyi Li, Wei Wang, Liya Ma, Fengjiao Chen, Xiaoyu Li, Xuezhi Cao, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Meituan(美团) Shanghai AI Lab(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Q-REAL数据集,用于细粒度评估AI生成图像的真实性和可信度,通过标注实体位置和设计判断问题,提升生成模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29967 2026-04-01 cs.CV 70%

Learning Structural-Functional Brain Representations through Multi-Scale Adaptive Graph Attention for Cognitive Insight

通过多尺度自适应图注意力学习结构-功能脑表示以获得认知洞察

Badhan Mazumder, Sir-Lord Wiafe, Aline Kotoski, Vince D. Calhoun, Dong Hye Ye

机构 * Georgia State University(佐治亚州立大学) Georgia Institute of Technology(佐治亚理工学院) Emory University(埃默里大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS)(三机构神经影像与数据科学转化研究中心 (TReNDS))

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MAGNet框架,通过多尺度自适应图注意力机制融合结构和功能脑网络,提升认知功能理解。

Comments Preprint version of the paper accepted to the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026). This is the author's accepted manuscript. The final published version will appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28651 2026-03-31 cs.AI 70%

Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning

不是搜索,而是扫描:在面向学术论文推理的扫描任务上基准测试大语言模型

Rongjin Li, Zichen Tang, Xianghe Wang, Xinyi Hu, Zhengyu Wang, Zhengyu Lu, Yiling Huang, Jiayuan Chen, Weisheng Tan, Jiacheng Liu, Zhongjun Yang, Haihong E

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出ScholScan基准,通过扫描式任务评估大语言模型在学术论文推理中的能力,发现检索增强生成方法在扫描任务上无显著提升,揭示了当前模型在该任务上的系统性缺陷。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24903 2026-03-31 cs.CV 70%

Self-Supervised Learning for Knee Osteoarthritis: Diagnostic Limitations and Prognostic Value of Hospital Data

自监督学习在膝骨关节炎中的应用:医院数据的诊断局限性与预后价值

Haresh Rengaraj Rajamohan, Yuxuan Chen, Kyunghyun Cho, Cem M. Deniz

机构 * Department of Physics, J.K. Institute of Science(物理系,J.K. 科学研究所) World Scientific University(世界科学大学) University of Intelligent Studies(智能研究大学) Center for Data Science, New York University(数据科学中心,纽约大学) Bernard and Irene Schwartz Center for Biomedical Imaging, New York University Langone Health(伯纳德和艾琳·施瓦茨生物医学成像中心,纽约大学朗格尼健康中心) Department of Radiology, New York University Langone Health(放射学系,纽约大学朗格尼健康中心) Perlmutter Cancer Center, New York University Langone Health(珀尔马特癌症中心,纽约大学朗格尼健康中心)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本研究评估自监督学习是否能提升膝骨关节炎的诊断与预后建模,发现医院数据在诊断分级中表现有限,但对预后预测有显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24569 2026-03-31 cs.CV 70%

POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan

POLY-SIM:基于缺失模态的多模态说话人识别挑战2026评估计划

Marta Moscati, Muhammad Saad Saeed, Marina Zanoni, Mubashir Noman, Rohan Kumar Das, Monorama Swain, Yufang Hou, Elisabeth Andre, Khalid Mahmood Malik, Markus Schedl, Shah Nawaz

机构 * Institute of Computational Perception, Johannes Kepler University Linz, Austria(林茨约翰·开普勒大学计算感知研究所) University of Michigan-Flint, USA(密歇根大学弗林特分校) Sapienza University of Rome, Italy(罗马大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Fortemedia Singapore, Singapore(新加坡Fortemedia公司) IT:U Interdisciplinary Transformation University Austria(奥地利跨学科转型大学) University of Augsburg, Germany(奥格斯堡大学) Human-centered AI Group, AI Lab, Linz Institute of Technology, Austria(奥地利林茨理工学院人工智能实验室人本人工智能组)

专题命中 多模态评测 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文针对多模态说话人识别中缺失模态和跨语言条件下的鲁棒性问题,提出POLY-SIM 2026挑战,设计标准化基准和评估框架以推动更实用的系统发展。

Comments Grand challenge at ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25791 2026-03-30 cs.CV 70%

ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions

ArtHOI:通过基础模型驯化实现单目4D手-物体交互重建

Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出ArtHOI框架,通过整合多基础模型先验知识,解决单目视频中手-物体交互的4D重建问题,引入自适应采样细化和多模态大语言模型引导对齐方法,构建两个新数据集并验证方法有效性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24198 2026-03-26 cs.CV 70%

RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution

RefReward-SR: 基于低分辨率参考的偏好对齐超分辨率奖励建模

Yushuai Song, Weize Quan, Weining Wang, Jiahui Sun, Jing Liu, Meng Li, Pengbin Yu, Zhentao Chen, Wei Shen, Lunxi Yuan, Dong-ming Yan

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) OPPO AI Center, OPPO Inc.(OPPO人工智能中心)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出RefReward-SR,通过低分辨率参考意识奖励模型实现偏好对齐的超分辨率,利用多模态大语言模型评估语义一致性,构建首个大规模低分辨率条件偏好数据集,实验表明其在人类判断对齐方面表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23883 2026-03-26 cs.CV 70%

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

BioVITA: 生物数据集、模型和基准用于视觉-文本-音频对齐

Risa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito, Hiroaki Santo, Fumio Okura

机构 * The University of Osaka(大阪大学) The University of Tokyo(东京大学) Institute of Science Tokyo(东京科学研究所) OMRON SINIC X

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出BioVITA框架,通过大规模数据集、模型和跨模态检索基准,实现生物物种的多模态对齐,提升生物多样性理解。

Comments CVPR 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22915 2026-03-25 cs.CV 70%

When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

当AVSR遇见视频会议:数据集、退化及性能崩溃的隐藏机制

Yihuan Huang, Jun Xue, Liu Jiajun, Daixian Li, Tong Zhang, Zhuolin Yi, Yanzhen Ren, Kai Li

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education(航空航天信息安全部与可信计算教育部重点实验室) School of Cyber Science and Engineering, Wuhan University(武汉大学网络安全科学与工程学院) Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文系统评估了主流视频会议平台上的最新AVSR模型,发现传输失真和人类超表达导致性能退化,通过构建MLD-VC数据集揭示语音增强算法引起分布偏移,细调模型可降低17.5%的CER。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21395 2026-03-25 cs.CV 70%

Momentum Memory for Knowledge Distillation in Computational Pathology

动量记忆用于计算病理学中的知识蒸馏

Yongxin Guo, Hao Lu, Onur C. Koyun, Zhengjie Zhu, Muhammet Fatih Demir, Metin Nafi Gurcan

机构 * Wake Forest University School of Medicine(威克森林大学医学院)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MoMKD框架,通过动量更新的记忆机制整合基因组和病理学信息,提升跨模态知识蒸馏性能,实验表明其在病理学任务中表现优异。

Comments Accepted by CVPR 2026. Code: https://github.com/CAIR-LAB-WFUSM/MoMKD

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11026 2026-03-24 cs.CV 70%

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

GIR-Bench:用于生成图像的多功能基准

Hongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu, Yuhang Yang, Xiaoshuang Huang, Jiayin Cai, Xiaolong Jiang, Yao Hu, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) Xiaohongshu Inc.(小红书公司)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出GIR-Bench,从理解-生成一致性、文本到图像生成和多步骤编辑三个角度评估统一模型,揭示其在复杂视觉任务中的表现与不足。

Comments ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18712 2026-03-20 cs.AI 70%

Accurate and Efficient Multi-Channel Time Series Forecasting via Sparse Attention Mechanism

通过稀疏注意力机制实现准确且高效的多通道时间序列预测

Lei Gao, Hengda Bao, Jingfei Fang, Guangzheng Wu, Weihua Zhou, Yun Zhou

机构 * Department of Machine Learning(机器学习系) SF Express(顺丰快递) Department of Management(管理系) Zhejiang University of Technology(浙江工业大学) Zhejiang University(浙江大学)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出Li-Net架构,通过稀疏Top-K Softmax注意力机制和多尺度投影框架,有效捕捉多通道时间序列的线性和非线性依赖,提升预测精度与效率。

Comments Accepted by ICDE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17746 2026-03-19 cs.CV 70%

Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation

概念到像素:无提示通用医学图像分割

Haoyun Chen, Fenghe Tang, Wenxin Ma, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei, Anhui 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu 215123, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出C2P框架,通过分离解剖学知识为几何和语义表示,利用多模态大语言模型生成语义令牌,并引入几何令牌约束,实现无提示的通用医学图像分割,实验表明其在多种模态和数据集上表现优异。

Comments 32 pages, code is available at: https://github.com/Yundi218/Concept-to-Pixel

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15967 2026-03-19 cs.CV 70%

A Comprehensive Benchmark of Histopathology Foundation Models for Kidney Digital Pathology Images

一种针对肾数字病理图像的病理基础模型综合基准测试

Harishwar Reddy Kasireddy, Patricio S. La Rosa, Akshita Gupta, Anindya S. Paul, Jamie L. Fermin, William L. Clapp, Meryl A. Waldman, Tarek M. El-Ashkar, Sanjay Jain, Luis Rodrigues, Kuang Yu Jen, Avi Z. Rosenberg, Michael T. Eadon, Jeffrey B. Hodgin, Pinaki Sarder

机构 * Department of Electrical and Computer Engineering, University of Florida(佛罗里达大学电气与计算机工程系) Seed Production Innovation, Crop Science Division, Bayer Company(拜耳公司种子生产创新部) Division of Medicine – Quantitative Health, University of Florida(佛罗里达大学医学部-定量健康分部) Department of Health Outcomes and Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系) Department of Pathology, Immunology and Laboratory Medicine, University of Florida College of Medicine(佛罗里达大学医学院病理学、免疫学与实验室医学系) Kidney Disease Branch, National Institute of Diabetes and Digestive and Kidney Diseases, National Institutes of Health(美国国立卫生研究院糖尿病、消化系统与肾病研究所肾病分支) Indiana University School of Medicine(印第安纳大学医学院) Departments of Medicine, Washington University School of Medicine(华盛顿大学医学院医学部) Universidade de Coimbra(科英布拉大学) Department of Pathology and Laboratory Medicine, University of California at Davis School of Medicine(加州大学戴维斯分校医学院病理学与实验室医学系) Department of Pathology, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院病理学系)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 本文评估了11种公开的病理基础模型在11个肾脏特定下游任务中的表现,揭示了其在肾病诊断中的适用性及改进方向。

Comments 31 Pages, 14 Tables, 12 figures, Co-correspondence to jhodgin@med.umich.edu and pinaki.sarder@ufl.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17413 2026-03-19 cs.CV 70%

Towards Motion-aware Referring Image Segmentation

面向运动感知的指认图像分割

Chaeyun Kim, Seunghoon Yi, Yejin Kim, Yohan Jo, Joonseok Lee

机构 * Seoul National University(首尔国立大学) AIM Intelligence(AIM智能)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出MRaCL方法,通过多模态径向对比学习提升运动相关查询的指认图像分割性能,并引入M-Bench基准测试运动主导的物体区分。

Comments Accepted at AISTATS 2026. * Equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15880 2026-03-18 cs.LG cs.AI 70%

Electrodermal Activity as a Unimodal Signal for Aerobic Exercise Detection in Wearable Sensors

电导活动作为可穿戴传感器中有氧运动检测的单一信号

Rena Mira Krishna, Ramya Sankar, Shadi Ghiasi

机构 * Odle Middle School(Odle中学) ImagineQ Labs(ImagineQ实验室) Cambridge Center for International(剑桥国际中心) bellevue, WA, USA(西雅图,华盛顿州,美国) Research, United Kingdom(英国研究)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 研究探讨了仅使用电导活动特征是否能可靠区分静息与持续有氧运动,采用留一被试法验证,发现EDA单独分类在主体独立评估中表现中等,相位时间动态和事件时间对分类分离有贡献。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15409 2026-03-17 cs.CL 70%

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

SEA-Vision:面向东南亚的多语言综合文档和场景文本理解基准

Pengfei Yue, Xingran Zhao, Juntao Chen, Peng Hou, Wang Longchao, Jianghang Lin, Shengchuan Zhang, Anxiang Zeng, Liujuan Cao

机构 * Xiamen University, China(厦门大学,中国) Shopee, China(Shopee,中国) Tongji University, China(同济大学,中国)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 SEA-Vision提出一个多语言基准,用于评估文档解析和文本中心视觉问答,涵盖11种东南亚语言,包含15234页文档和7496对问答对,揭示多语言文档理解的差距。

Comments Accepted By CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12266 2026-03-13 cs.CV 70%

MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning

MM-CondChain: 一种可编程验证的视觉基础深度组合推理基准

Haozhan Shen, Shilin Yan, Hongwei Xue, Shuaiqi Lu, Xiaojun Tang, Guannan Zhang, Tiancheng Zhao, Jianwei Yin

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 MM-CondChain提出了一种可编程验证的视觉基础深度组合推理基准,通过多层推理链评估多模态大语言模型在复杂视觉任务中的表现,实验表明深度组合推理仍是重大挑战。

Comments Project Page: https://accio-lab.github.io/MM-CondChain

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09774 2026-03-11 cs.AI 70%

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

World2Mind: 用于基础模型的方位认知工具包

Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) Huawei Noah’s Ark Lab(华为诺亚方舟实验室) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学人工智能院计算机科学与技术系、清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 World2Mind通过构建空间认知地图和三阶段推理链,提升基础模型的空间推理能力,使纯文本模型也能实现复杂3D空间推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06700 2026-03-10 cs.CV 70%

SIQA: Toward Reliable Scientific Image Quality Assessment

SIQA:迈向可靠的科学图像质量评估

Wenzhe Li, Liang Chen, Junying Wang, Yijing Guo, Ye Shen, Farong Wen, Chunyi Li, Zicheng Zhang, Guangtao Zhai

机构 * TongJi University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 SIQA提出了一种多维框架,通过SIQA-U和SIQA-S评估科学图像的科学正确性和感知清晰度,揭示模型在评分一致性与科学理解上的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21950 2026-03-03 cs.CV 70%

Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach

为MLLMs定制视觉情绪评估:一种开放词汇、多维且可扩展的方法

Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理学与认知科学系) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出一种开放词汇、多维且可扩展的方法,用于定制MLLMs的视觉情绪评估,通过情绪陈述判断任务和自动化流程,评估现有模型在情绪识别和上下文判断中的表现,并揭示其局限性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13671 2026-03-03 cs.CV cs.LG 70%

FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization

FiLo:通过细粒度描述和高质量定位实现零样本异常检测

Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Hao Li, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所基础模型研究中心) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) Objecteye Inc., Beijing, China(Objecteye公司) Central South University, Hunan, China(中南大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 FiLo通过细粒度描述和高质量定位实现零样本异常检测,提升检测和定位性能。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23730 2026-03-02 cs.AI 70%

Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off

解锁认知能力并分析感知-逻辑权衡

Longyin Zhang, Shuo Sun, Yingxu He, Won Cheng Yi Lewis, Muhammad Huzaifah Bin Md Shahrin, Hardik Bhupendra Sailor, Heng Meng Jeremy Wong, Tarun Kumar Vangani, Yi Ma, Qiongqiong Wang, Minh Duc Pham, Ridong Jiang, Jingtao Li, Jingyi Liao, Zhuohan Liu, Yanfeng Lu, Manas Gupta, Ai Ti Aw

机构 * Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究 institute(I 2 R),A*STAR,新加坡)

专题命中 多模态评测 :multimodal(abstract);audio-visual(abstract);分类 cs.AI

AI总结 MERaLiON2-Omni(Alpha)通过解耦感知与推理能力,针对东南亚地区提出多语言全方位感知模型,揭示感知与逻辑之间的效率-稳定性权衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22033 2026-02-26 cs.CV 70%

RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking

RT-RMOT:一种用于RGB-热成像参照多目标跟踪的数据集和框架

Yanqiu Yu, Zhifan Jin, Sijia Chen, Tongfei Chu, En Yu, Liman Liu, Wenbing Tao

机构 * Huazhong University of Science and Technology(华中科技大学) South-Central Minzu University(西南民族大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 RT-RMOT提出一种融合RGB和热成像特征的多目标跟踪框架,通过改进的RL策略优化提升全天候跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19828 2026-02-24 cs.CV 70%

TextShield-R1: Reinforced Reasoning for Tampered Text Detection

TextShield-R1: 用于篡改文本检测的强化推理

Chenfan Qu, Yiwu Zhong, Jian Liu, Xuekang Zhu, Bohan Yu, Lianwen Jin

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 TextShield-R1通过强化学习和OCR校正技术,提升篡改文本检测的准确性和鲁棒性,解决现有方法在微特征识别和跨语言评估中的不足。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26961 2026-02-24 cs.CV eess.IV 70%

SYNAPSE-Net: A Unified Framework with Lesion-Aware Hierarchical Gating for Robust Segmentation of Heterogeneous Brain Lesions

SYNAPSE-Net:一种具有病变感知分层门控的统一框架,用于鲁棒分割异质性脑病变

Md. Mehedi Hassan, Shafqat Alam, Shahriar Ahmed Seam, Maruf Ahmed

机构 * Department of Biomedical Engineering, Bangladesh University of Engineering and Technology(生物医学工程系,孟加拉国工程与技术大学) Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology(计算机科学与工程系,孟加拉国工程与技术大学) Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SYNAPSE-Net通过统一框架和病变感知门控,实现鲁棒的多病理脑病变分割,提升分割精度与稳定性。

Comments 18 pages, 10 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17096 2026-02-20 cs.AI 70%

Agentic Wireless Communication for 6G: Intent-Aware and Continuously Evolving Physical-Layer Intelligence

面向6G的代理无线通信:意图感知与持续演化的物理层智能

Zhaoyang Li, Xingzhi Jin, Junyu Pan, Qianqian Yang, Zhiguo Shi

机构 * College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于LLM的代理AI用于6G物理层,通过意图感知和持续演进实现自主通信决策。

详情

展开后加载摘要…

URL PDF HTML 收藏