arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-03 至 2026-03-03 共收录 206 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 26 篇

2506.10941 2026-03-03 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

VINCIE: Unlocking In-context Image Editing from Video

VINCIE:从视频中解锁上下文图像编辑

Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang

机构 * National University of Singapore(国立新加坡大学) ByteDance Seed(字节跳动种子)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 VINCIE通过视频直接训练模型,实现了强大的上下文图像编辑能力,并在多轮图像编辑基准中取得最佳成绩。

Comments ICLR 2026 Camera-ready. Project page: https://vincie2025.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00159 2026-03-03 cs.CV cs.AI cs.MM cs.SD 67%

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

FlowPortrait:基于强化学习的音频驱动肖像视频生成

Weiting Tan, Andy T. Liu, Ming Tu, Xinghua Qu, Philipp Koehn, Lu Lu

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 FlowPortrait通过强化学习框架结合多模态模型和人类对齐评估系统,提升音频驱动肖像视频生成的质量和表现力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14341 2026-03-03 cs.CV cs.AI cs.CY cs.LG 62%

Towards Transferable Defense Against Malicious Image Edits

面向恶意图像编辑的可迁移防御

Jie Zhang, Shuai Dong, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of China Academy of Sciences(中国科学院大学) School of Computer Science, China University of Geosciences(中国地质大学(武汉)计算机学院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 TDAE通过双模优化提升图像对恶意编辑的免疫性,实现跨模型的可迁移防御。

Comments 14 pages, 5 figures, accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02253 2026-03-03 cs.CV cs.AI cs.LG 62%

DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing

DragFlow: 通过基于区域的监督释放DiT先验以实现拖拽编辑

Zihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang, Zhuming Lian, Xinlei Yu, Adams Wai-Kin Kong

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(国立新加坡大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 DragFlow通过基于区域的监督利用FLUX先验,改进基于拖拽的图像编辑效果,实现对点式和区域式基线的超越。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16145 2026-03-03 cs.AI cs.CL 62%

SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting

SpiroLLM:通过临床验证在COPD报告中微调预训练大语言模型以理解肺功能时间序列

Shuhao Mei, Yongchao Long, Xiaoyu Xiao, Shan Cao, Xiaobo Han, Shijia Geng, Jinbo Sun, Yuxi Zhou, Shenda Hong

机构 * Guangzhou Institute of Technology, Xidian University, Xi’an, China(广州科技研究院,西安电子科技大学,中国) Department of Computer Science, Tianjin University of Technology, Tianjin, China(天津理工大学计算机学院,天津,中国) Department of Respiratory, The Second Hospital of Tianjin Medical University, China(天津医科大学第二医院呼吸科,中国) College of Pulmonary and Critical Care Medicine, Chinese PLA General Hospital, Beijing, China(中国人民解放军总医院呼吸与危重症医学科,北京,中国) HeartVoice Medical Technology, Hefei, China(合肥心声医疗技术有限公司,中国) School of Life Science and Technology, Xidian University, Xi’an, China(西安电子科技大学生命科学与技术学院,中国) National Institute of Health Data Science, Peking University, Beijing, China(北京大学国家健康数据科学研究院,北京,中国) Institute for Artificial Intelligence, Peking University, Beijing, China(北京大学人工智能研究院,北京,中国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 SpiroLLM通过融合生理信号与大语言模型,实现对肺功能时间序列的解读,并在COPD诊断中展现出高准确性和稳健性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12734 2026-03-03 cs.SD cs.AI cs.GR cs.HC eess.AS 62%

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

SounDiT:基于地理情境的声音景观到景观生成

Junbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang, Teng Fei, Qixing Huang, Bing Zhou, Zhengzhong Tu, Shan Ye, Yuhao Kang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Tennessee, Knoxville(田纳西大学基洛纳分校) University of South Carolina(南卡罗来纳大学) Arizona State University(亚利桑那州立大学) University of Canterbury(坎特伯雷大学) Texas A&M University(德克萨斯A&M大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI、eess.AS

AI总结 SounDiT通过结合环境声音景观和地理情境条件,生成地理上一致的景观图像,并引入Place Similarity Score评估生成一致性。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00155 2026-03-03 cs.CV cs.AI cs.IR 62%

EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection

EfficientPosterGen: 通过令牌压缩和准确违规检测的语义感知高效海报生成

Wenxin Tang, Jingyu Xiao, Yanpei Gong, Fengyuan Ran, Tongchuan Xia, Junliang Liu, Man Ho Lam, Wenxuan Wang, Michael R. Lyu

机构 * Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Harbin Institute of Technology(哈尔滨工业大学) Wuhan University(武汉大学) Beijing University of Posts and Telecommunications(北京邮电大学) Dalian Maritime University(大连海事大学) Renmin University of China(中国人民大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 EfficientPosterGen通过语义感知检索、视觉上下文压缩和无代理布局检测技术,实现高效且可靠的学术海报自动生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02138 2026-03-03 cs.CV 57%

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

OmniLottie:通过参数化Lottie令牌生成向量动画

Yiying Yang, Wei Cheng, Sijin Chen, Honghao Fu, Xianfang Zeng, Yujun Cai, Gang Yu, Xingjun Ma

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 OmniLottie通过参数化Lottie令牌生成高质量向量动画,结合多模态指令与大规模数据集提升动画生成能力。

Comments Accepted by CVPR 2026. Project Page: https://openvglab.github.io/OmniLottie/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01552 2026-03-03 cs.CV 57%

Align-cDAE: Alzheimer's Disease Progression Modeling with Attention-Aligned Conditional Diffusion Auto-Encoder

Align-cDAE: 利用注意力对齐的条件扩散自编码器进行阿尔茨海默病进展建模

Ayantika Das, Keerthi Ram, Mohanasankar Sivaprakasam

机构 * Department of Electrical Engineering, Indian Institute of Technology Madras, Chennai, India(电子工程系,印度理工学院马德拉斯,钦奈,印度) Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India(苏达·戈帕拉克里希南脑中心,印度理工学院马德拉斯,钦奈,印度) Department of Electrical Engineering, Indian Institute of Technology Madras Chennai, India(电子工程系,印度理工学院马德拉斯钦奈,印度)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 Align-cDAE通过引入注意力对齐和结构化潜在空间,提升扩散自编码器在阿尔茨海默病进展建模中的精度和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01068 2026-03-03 cs.CV cs.LG 57%

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

LLaDA-o:一种高效且长度自适应的多模态扩散模型

Zebin You, Xiaolu Zhang, Jun Zhou, Chongxuan Li, Ji-Rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China.(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models(北京大型模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 LLaDA-o通过混合扩散框架和数据驱动的长度适应策略,实现了高效且灵活的多模态扩散建模,展示了在文本到图像生成任务中的卓越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11758 2026-03-03 q-bio.QM cs.AI 57%

Protein Structure Tokenization via Geometric Byte Pair Encoding

通过几何字对编码进行蛋白质结构分词

Michael Sun, Weize Yuan, Gang Liu, Wojciech Matusik, Marinka Zitnik

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Harvard Medical School(哈佛医学院) Apple(苹果公司) Notre Dame(诺特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 GeoBPE通过几何字对编码实现蛋白质结构分词,提供压缩、数据效率和泛化能力,支持多架构应用并增强功能解释性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00521 2026-03-03 cs.LG cs.AI 57%

Phys-Diff: A Physics-Inspired Latent Diffusion Model for Tropical Cyclone Forecasting

Phys-Diff:一种基于物理的潜在扩散模型用于热带气旋预测

Lei Liu, Xiaoning Yu, Kang Chen, Jiahui Huang, Tengyuan Liu, Hongwei Zhao, Bin Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 Phys-Diff通过结合物理启发的归纳偏置和多模态数据整合,提升热带气旋预测的物理一致性与性能

Comments 5 pages, 4 figures. Accepted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00266 2026-03-03 cs.CV 57%

Adversarial Patch Generation for Visual-Infrared Dense Prediction Tasks via Joint Position-Color Optimization

为视觉-红外密集预测任务生成对抗性补丁的联合位置-颜色优化

He Li, Wenyue He, Weihang Kong, Xingchen Zhang

机构 * Yanshan University(燕山大学) University of Exeter(埃克塞特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出AP-PCO框架,通过联合优化位置和颜色生成对抗性补丁,提升视觉-红外密集预测任务中的攻击性能和隐蔽性。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10980 2026-03-03 cs.CV 57%

TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation

TrueSkin: 向公平和准确的皮肤色调识别与生成迈进

Haoming Lu

机构 * Topaz Labs(Topaz实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 TrueSkin通过系统化的皮肤色调数据集,提升模型在公平性和准确性上的表现,验证了其在识别和生成任务中的有效性。

Comments The dataset is available for download at https://drive.google.com/file/d/1_ndw5uyY4h4DLL5iGTL4bVDKdE_g_H4B/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00116 2026-03-03 cs.CV 57%

VoxelDiffusionCut: Non-destructive Internal-part Extraction via Iterative Cutting and Structure Estimation

VoxelDiffusionCut:通过迭代切割和结构估计实现非破坏性内部部件提取

Takumi Hachimine, Yuhwan Kwon, Cheng-Yu Kuo, Tomoya Yamanokuchi, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,奈良科学技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 VoxelDiffusionCut通过迭代切割和结构估计方法,利用扩散模型估计体素结构以实现非破坏性内部部件提取。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 44 篇

2506.09427 2026-03-03 cs.CV cs.AI 86%

A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation

一种高质量数据集和可靠的评估方法用于交错图像-文本生成

Yukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li, Jiaxin Ai, Fanrui Zhang, Yifan Chang, Sizhuo Zhou, Shenglin Zhang, Yu Dai, Kaipeng Zhang

机构 * Nankai University(南开大学) Shanghai Innovation Institute(上海创新研究院) Shanda AI Research(尚德人工智能研究院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :image-text(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 InterSyn数据集通过高质量、大规模和多样化指导设计,提升LMMs在图像-文本生成任务中的性能,并提出SynJudge评估方法,全面评估内容和跨模态交互质量。

Comments Accepted in ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06749 2026-03-03 cs.CV cs.AI cs.CL cs.LG 85%

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Vision-R1: 促进多模态大语言模型推理能力的激励方法

Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Xu Tang, Yao Hu, Shaohui Lin

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Xiaohongshu Inc.(小红书公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Vision-R1通过构建高质量多模态CoT数据集和渐进性思维抑制训练策略,提升多模态推理能力,在数学推理基准中取得显著成绩。

Comments Accepted to ICLR 2026. Code is available at https://github.com/Osilly/Vision-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01475 2026-03-03 cs.CV 83%

WildCross: A Cross-Modal Large Scale Benchmark for Place Recognition and Metric Depth Estimation in Natural Environments

WildCross: 一种用于自然环境中地方识别和度量深度估计的跨模态大规模基准

Joshua Knights, Joseph Reid, Kaushik Roy, David Hall, Mark Cox, Peyman Moghadam

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 WildCross是一个用于自然环境中地方识别和度量深度估计的跨模态大规模基准,通过提供大规模数据集和实验验证,推动多模态机器人感知技术的发展。

Comments IEEE International Conference on Robotics & Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17237 2026-03-03 cs.CV 83%

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

Grounding-IQA: 多模态语言模型的接地图像质量评估

Zheng Chen, Xun Zhang, Wenbo Li, Renjing Pei, Fenglong Song, Xiongkuo Min, Xiaohong Liu, Xin Yuan, Yong Guo, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Joy Future Academy(京东探索研究院) Huawei Noah’s Ark Lab(华为诺亚实验室) Westlake University(西湖大学) Huawei(华为)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出grounding-IQA任务范式,通过结合多模态指称与图像质量评估,实现更细粒度的图像质量感知。

Comments Accepted to ICLR 2026. Code is available at: https://github.com/zhengchen1999/Grounding-IQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00873 2026-03-03 cs.AI 83%

MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains

MC-Search:基于结构化长推理链评估和增强多模态代理搜索

Xuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai, Jiaru Zou, Ting-Wei Li, Hanghang Tong, Yada Zhu, Hendrik Hamann, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Meta IBM Research(IBM研究院) Stony Brook University(石溪大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MC-Search是首个针对代理MM-RAG的基准,通过结构化长推理链评估和增强多模态代理搜索,揭示了系统性问题并改进了模型的规划和检索能力。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00971 2026-03-03 cs.CV 83%

Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning

揭示认知罗盘:基于理论of-Mind的多模态情感推理

Meng Luo, Bobo Li, Shanqing Xu, Shize Zhang, Qiuchan Chen, Menglu Han, Wenhao Chen, Yanxiang Huang, Hao Fei, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HitEmotion基准和TMPO方法,通过理论of-Mind引导多模态情感推理,提升模型情感理解和认知能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00108 2026-03-03 cs.RO cs.AI 83%

SurgFusion-Net: Diversified Adaptive Multimodal Fusion Network for Surgical Skill Assessment

SurgFusion-Net:用于外科技能评估的多样化自适应多模态融合网络

Runlong He, Freweini M. Tesfai, Matthew W. E. Boal, Nazir Sirajudeen, Dimitrios Anastasiou, Jialang Xu, Mobarak I. Hoque, Philip J. Edwards, John D. Kelly, Ashwin Sridhar, Abdolrahim Kadkhodamohammadi, Dhivya Chandrasekaran, Matthew J. Clarkson, Danail Stoyanov, Nader Francis, Evangelos B. Mazomenos

机构 * UCL Hawkes Institute and the Department of Medical Physics & Biomedical Engineering, UCL(UCL哈维斯研究所及UCL医学物理与生物医学工程系) UCL Hawkes Institute and the Department of Computer Science, UCL(UCL哈维斯研究所及UCL计算机科学系) UCL Hawkes Institute and the Division of Informatics, Imaging & Data Sciences, The University of Manchester(UCL哈维斯研究所及信息学、成像与数据科学系,曼彻斯特大学) UCL Hospitals NHS Foundation Trust(UCL医院 NHS基金会信托) Griffin Institute, Northwick Park and St Mark’s Hospital(格里芬研究所,北wick公园及圣马可医院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 SurgFusion-Net通过引入DRA策略和两个新的临床数据集,提升了多模态外科技能评估的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02024 2026-03-03 cs.CL cs.AI cs.CV 82%

MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

MMR-Life: 组合真实场景以进行多模态多图像推理

Jiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所认知与决策智能复杂系统重点实验室) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学交叉学科学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MMR-Life是一个评估多模态多图像推理能力的综合基准,通过现实场景中的多图像信息整合和多种推理类型测试,揭示了现有模型在复杂推理任务中的挑战和性能差异。

Comments Accepted by ICLR 2026, 78 pages, 60 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02185 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models

视觉-深度研究基准:重新思考多模态大语言模型的视觉和文本搜索

Yu Zeng, Wenxuan Huang, Zhen Fang, Shuang Chen, Yufan Shen, Yishuo Cai, Xiaoman Wang, Zhenfei Yin, Lin Chen, Zehui Chen, Shiting Huang, Yiming Zhao, Xu Tang, Yao Hu, Philip Torr, Wanli Ouyang, Shaosheng Cao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出Vision-DeepResearch基准,通过精心设计的问题评估多模态大语言模型的视觉与文本搜索能力,并引入多轮裁剪搜索流程提升实际检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02200 2026-03-03 cs.CV cs.AI cs.LG 81%

Adaptive Confidence Regularization for Multimodal Failure Detection

多模态故障检测的自适应置信度正则化

Moru Liu, Hao Dong, Olga Fink, Mario Trapp

机构 * Technical University of Munich(慕尼黑技术大学) ETH Zürich(苏黎世联邦理工学院) EPFL(苏黎世联邦理工学院) Fraunhofer IKS(弗劳恩霍夫研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出自适应置信度正则化方法,通过检测多模态预测中的置信度退化,提升故障识别的可靠性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00041 2026-03-03 cs.CV cs.AI 81%

Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness

文化在框架中:C$^3$B作为基于漫画的多模态文化意识基准

Yuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen, Xuefeng Bai, Muyun Yang, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 C$^3$B是一个基于漫画的多模态文化意识基准,旨在评估和提升多语言、多任务和多文化环境下MLLMs的文化理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00115 2026-03-03 physics.soc-ph cs.AI cs.CV 81%

Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment

多模态模块化思维链在能源性能证书评估中的应用

Zhen Peng, Peter J. Bentley

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种基于多模态模块化思维链的低成本EPC预评估方法,通过结构化提示实现对EPC评分的序数结构捕捉,实验表明其在数据稀缺环境下具有显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06499 2026-03-03 cs.CV 79%

SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

SportR:多模态大语言模型在体育中的推理基准

Haotian Xia, Haonan Ge, Junbo Zou, Hyun Woo Choi, Xuebin Zhang, Danny Suradja, Botao Rui, Ethan Tran, Wendy Jin, Zhen Ye, Xiyang Lin, Christopher Lai, Shengjie Zhang, Junwen Miao, Shichao Chen, Rhys Tracy, Vicente Ordonez, Weining Shen, Hanjie Chen

机构 * Department of Computer Science, Rice University(Rice大学计算机科学系) Ken Kennedy Institute, Rice University(Rice大学肯尼迪研究所) Department of Statistics, University of California, Irvine(伊利诺伊大学欧文分校统计系) College of Sciences, Georgia Institute of Technology(佐治亚理工学院科学学院) Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系) Department of Computer Science, University of California, Santa Barbara(加州大学圣芭芭拉分校计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 SportR是一个多体育大规模基准,旨在训练和评估多模态大语言模型在体育推理中的能力,通过精细的视觉感知和规则推理任务提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01557 2026-03-03 cs.AI 79%

Benchmarking LLM Summaries of Multimodal Clinical Time Series for Remote Monitoring

对多模态临床时间序列远程监测的LLM摘要进行基准测试

Aditya Shukla, Yining Yuan, Ben Tamo, Yifei Wang, Micky Nnamdi, Shaun Tan, Jieru Li, Benoit Marteau, Brad Willingham, May Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Shepherd Center(Shepherd中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于事件的评估框架,评估多模态临床时间序列摘要的可靠性,发现视觉方法在事件对齐上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01547 2026-03-03 cs.CV 79%

PathMoE: Interpretable Multimodal Interaction Experts for Pediatric Brain Tumor Classification

PathMoE:用于儿童脑肿瘤分类的可解释多模态交互专家

Jian Yu, Joakim Nguyen, Jinrui Fang, Awais Naeem, Zeyuan Cao, Sanjay Krishnan, Nicholas Konz, Tianlong Chen, Chandra Krishnan, Hairong Wang, Edward Castillo, Ying Ding, Ankita Shukla

机构 * University of Texas(德克萨斯大学) Dell Children’s Medical Center(德尔儿童医学中心) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Nevada, Reno(内华达大学里诺分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 PathMoE通过整合多模态信息提升儿童脑肿瘤分类性能,揭示了不同模态间的交互作用。

详情

展开后加载摘要…

URL PDF HTML 收藏