arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-19 至 2025-12-19 共收录 51 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2406.09121 2025-12-19 cs.CV 79%

MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models

MMRel:多模态大语言模型中关系理解的基准测试

Jiahao Nie, Gongjie Zhang, Wenbin An, Yun Xing, Yap-Peng Tan, Alex C. Kot, Shijian Lu

机构 * Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(南洋理工大学跨学科研究生项目) Alibaba DAMO Academy, Singapore(阿里巴巴达摩院) Xi’an Jiaotong University, China(西安交通大学) Nanyang Technological University, Singapore(南洋理工大学) VinUniversity, Vietnam(文莱大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 MMRel是一个用于评估和提升多模态大语言模型关系理解能力的基准测试,包含大规模高质量的关系数据和对抗性案例,通过实验验证其在提升模型关系理解方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16202 2025-12-19 cs.CV cs.AI 62%

Open Ad-hoc Categorization with Contextualized Feature Learning

开放性即需分类与上下文化特征学习

Zilin Wang, Sangwoo Mo, Stella X. Yu, Sima Behpour, Liu Ren

机构 * University of Michigan(密歇根大学) UC Berkeley(加州大学伯克利分校) Bosch Center for AI(博世人工智能中心)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 OAK通过引入上下文标记和结合CLIP与GCD目标,实现了开放性即需分类的高准确率和可解释性。

Comments 26 pages, 17 figures

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13089 2025-12-19 cs.CV cs.AI 62%

UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era

UniVCD:面向开放词汇时代的无监督变化检测新方法

Ziqiang Zhu, Bowei Yang

机构 * School of Aeronautics and Astronautics, Zhejiang University(航空宇航学院,浙江大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 UniVCD是一种基于冻结视觉基础模型的无监督开放词汇变化检测方法,通过轻量级多模态对齐实现高分辨率语义感知变化检测。

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16561 2025-12-19 cs.CV 57%

N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models

N3D-VLM:原生3D接地使视觉-语言模型在3D场景中实现精确的空间推理

Yuxin Wang, Lei Ke, Boqiang Zhang, Tianyuan Qu, Hanxun Yu, Zhenpeng Huang, Meng Yu, Dan Xu, Dong Yu

机构 * HKUST(香港科技大学) Tencent AI Lab(腾讯AI实验室) CUHK(香港中文大学) ZJU(浙江大学) NJU(南京大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 N3D-VLM通过原生3D感知能力提升视觉-语言模型在3D场景中的空间推理精度与解释性。

Comments Project Page: https://n3d-vlm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15940 2025-12-19 cs.CV cs.RO 57%

R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space

R4:在4D时空空间中为视觉语言模型引入检索增强推理

Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Esslingen University of Applied Sciences(埃斯林根应用科学大学) Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业协会) University of Michigan(密歇根大学) Voxel51 Inc.(Voxel51公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 R4通过构建4D时空知识库,使视觉语言模型具备结构化终身记忆,实现无需训练的检索增强推理,提升动态环境中的具身推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15233 2025-12-19 cs.CV 57%

Null-LoRA: Low-Rank Adaptation on Null Space

Null-LoRA: 在空域上进行低秩适应

Yi Zhang, Yulei Kang, Haoxuan Chen, Jinxuan Li, Jian-Fang Hu

机构 * School of Computer Science(计算机科学学院) Engineering, Sun Yat-sen University, Guangdong, China(工程学院,中山大学,广东,中国)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 Null-LoRA通过在空域上进行低秩适应,提升参数效率和模型效果,在图像-文本检索和视觉问答任务中取得最优表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12977 2025-12-19 cs.CV 57%

VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference

VLCache:计算2%的视觉令牌并重用98%用于视觉-语言推断

Shengling Qin, Hao Yu, Chenxin Wu, Zheng Li, Yizhong Cao, Zhengyang Zhuge, Yuxin Zhou, Wentao Yao, Yi Zhang, Zhengheng Wang, Shuai Bai, Jianwei Zhang, Junyang Lin

机构 * Qwen Team, Alibaba Inc.(通义实验室,阿里巴巴集团)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 VLCache通过重用98%的缓存减少计算量,实现高效视觉-语言推理,精度与全重新计算相当,加速1.2x至16倍。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2512.01185 2025-12-19 cs.CR 85%

DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks

DefenSee:从视觉和文本中解构威胁——一种多视图防御管道用于多模态对抗突破

Zihao Wang, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 音频语音多模态 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract)

AI总结 DefenSee通过图像变体转录和跨模态一致性检查,提供一种多模态防御方法,有效降低多模态对抗攻击的成功率,提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16250 2025-12-19 cs.AI cs.MA 83%

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

AMUSE:面向代理多说话者理解的音频-视觉基准与对齐框架

Sanjoy Chowdhury, Karren D. Yang, Xudong Liu, Fartash Faghri, Pavan Kumar Anasosalu Vasu, Oncel Tuzel, Dinesh Manocha, Chun-Liang Li, Raviteja Vemulapalli

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Apple(苹果公司)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 AMUSE提出一个面向多说话人理解的音频-视觉基准与对齐框架RAFT,通过代理推理提升多模态模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14491 2025-12-19 cs.LG cs.MM 79%

Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review

多模态方法用于分析学习与训练环境:系统文献综述

Clayton Cohn, Eduardo Davalos, Caleb Vatral, Joyce Horn Fonteles, Hanchen David Wang, Austin Coursey, Surya Rayala, Ashwin T S, Meiyi Ma, Gautam Biswas

机构 * Vanderbilt University(范德比大学) Tennessee State University(田纳西州立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

AI总结 本文综述了多模态方法在学习与训练环境中的应用,提出分类法和框架,揭示了多模态整合对行为分析的价值及现存挑战。

Comments Submitted to ACM Computing Surveys. Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07741 2025-12-19 cs.LG cs.SD 78%

A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data

一种多模态贝叶斯网络用于从语音和语音数据中预测症状层面的抑郁和焦虑

Agnes Norbury, George Fairs, Alexandra L. Georgescu, Matthew M. Nour, Emilia Molimpakis, Stefano Goria

机构 * thymia Limited(thymia有限公司) Institute of Psychiatry, Psychology & Neuroscience, King’s College London(心理学与神经科学研究院,伦敦国王学院) Department of Psychiatry, University of Oxford(牛津大学精神病学系) Max Planck UCL Centre for Computational Psychiatry and Ageing, University College London(Max Planck大学学院计算精神病学与衰老中心,伦敦大学学院)

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 本文提出了一种多模态贝叶斯网络模型,用于从语音和语音数据中预测抑郁和焦虑症状,通过评估模型性能和公平性,展示了其在临床应用中的潜力。

Journal ref Scientific Reports (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12424 2025-12-19 cs.LG cs.AI cs.IR 70%

Multi-Modality Collaborative Learning for Sentiment Analysis

多模态协作学习用于情感分析

Shanmin Wang, Chengguang Liu, Qingshan Liu

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出多模态协作学习框架,通过解耦模块和策略模型提升跨模态情感特征学习,实验证明在四个数据库上性能显著提升。

Comments The method has flaws, especially with the decoupling module. During the decoupling process, the heterogeneity of the three modal data and the differences in distribution were not taken into account

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2512.16023 2025-12-19 cs.CV 79%

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

CoVAR: 通过多模态扩散生成视频与动作用于机器人操作

Liudi Yang, Yang Bai, George Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen, Ziyuan Liu, Abhinav Valada

机构 * University of Freiburg(弗赖堡大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Technical University of Munich(慕尼黑技术大学) Huawei Heisenberg Research Center (Munich)(华为海森堡研究中心)

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV

AI总结 CoVAR通过多模态扩散模型生成视频与动作,解决机器人操作中动作标注不足的问题,提升视频生成质量与动作精度。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18951 2025-12-19 cs.CV 79%

Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition

感知、对话,然后适应:面向开放世界视频识别的多模态基础模型知识迁移

Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PCA框架,通过感知、对话和适应三个阶段,利用多模态知识提升开放世界视频识别的性能。

Comments 35 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16842 2025-12-19 cs.CV cs.AI cs.RO 73%

OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction

OPENTOUCH:将全手触觉带入现实世界交互

Yuxin Ray Song, Jinzhou Li, Rao Fu, Devin Murphy, Kaichen Zhou, Rishi Shiv, Yaqi Li, Haoyu Xiong, Crystal Elaine Owens, Yilun Du, Yiyue Luo, Xianyi Cheng, Antonio Torralba, Wojciech Matusik, Paul Pu Liang

机构 * MIT(麻省理工学院) Duke University(杜克大学) Brown University(布朗大学) University of Washington(华盛顿大学) Harvard University(哈佛大学)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 OpenTouch是首个现实场景的第一人称全手触觉数据集,通过同步视频-触觉-姿态数据和详细注释,推动多模态感知和机器人操作研究。

Comments https://opentouch-tactile.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16924 2025-12-19 cs.CV 57%

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

世界是你的画布:通过参考图像、轨迹和文本绘画可提示事件

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Yue Yu, Yihao Meng, Wen Wang, Ka Leong Cheng, Shuailei Ma, Qingyan Bai, Yixuan Li, Cheng Chen, Yanhong Zeng, Xing Zhu, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) ZJU(浙江大学) NEU(南京大学) CUHK(香港中文大学) NTU(南洋理工大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 WorldCanvas通过结合文本、轨迹和参考图像,实现可提示的多代理交互事件生成,提升世界模型的交互性和可控性。

Comments Project page and code: https://worldcanvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16461 2025-12-19 cs.CV cs.RO 57%

SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning

SNOW:基于世界知识的时空场景理解用于开放世界具身推理

Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Esslingen University of Applied Sciences(埃斯林根应用科学大学) Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业联合会) University of Michigan(密歇根大学) Voxel51 Inc.(Voxel51公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SNOW通过整合视觉语言模型与点云几何,实现统一的4D场景理解,提升开放世界具身推理的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2512.16802 2025-12-19 cs.CL 79%

Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology

多模态检索增强生成在生物医学领域中的增强策略探索:一项评估糖生物学问答的案例研究

Primož Kocbek, Azra Frkatović-Hodžić, Dora Lalić, Vivian Hui, Gordan Lauc, Gregor Štiglic

机构 * University of Maribor, Faculty of Health Sciences(莫拉维亚大学健康科学学院) University of Ljubljana, Medical Factory(卢布尔雅那大学医疗工厂) Genos Ltd(基因公司) Center for Smart Health, School of Nursing The Hong Kong Polytechnic University(智能健康中心护理学院香港理工大学) University of Zagreb, Faculty of Pharmacy(扎格雷布大学药学院) Usher Institute University of Edinburgh(埃德蒙顿大学usher研究所)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文研究了多模态检索增强生成在生物医学领域中的增强策略,通过实验发现多模态转换和视觉检索在不同模型中均能提升问答准确率。

Comments Will be published in IEEE BigData 2025 proceedings. Contains 10 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08976 2025-12-19 cs.LG cs.AI 57%

Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs

peek-a-boo推理:多模态大语言模型中的对比区域遮挡

Isha Chaturvedi, Anjana Nair, Yushen Li, Adhitya Rajendra Kumar, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma

机构 * Algoverse AI Research(Algoverse AI研究机构) Princeton University(普林斯顿大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 对比区域遮挡通过遮挡视觉区域并对比推理轨迹,揭示多模态大语言模型在推理过程中的依赖模式和失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05020 2025-12-19 cs.GR cs.CV 57%

DAFM: Dynamic Adaptive Fusion for Multi-Model Collaboration in Composed Image Retrieval

DAFM:动态自适应融合用于复合图像检索中的多模型协作

Yawei Cai, Jiapeng Mi, Nan Ji, Haotian Rong, Yawei Zhang, Zhangti Li, Wenbin Guo, Rensong Xie

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 DAFM通过动态自适应融合多模型优势,提升复合图像检索的准确性和鲁棒性。

Comments We discovered an error that affects the main conclusions, so we decided to withdraw the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01146 2025-12-19 q-bio.OT 50%

Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications

生物医学中的检索增强生成:技术、数据集和临床应用的综述

Jiawei He, Boya Zhang, Hossein Rouhizadeh, Yingjian Chen, Rui Yang, Jin Lu, Xudong Chen, Nan Liu, Douglas Teodoro

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文综述了生物医学中检索增强生成技术的发展,探讨了其在临床应用中的挑战与未来发展方向。

Comments 49 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2512.13107 2025-12-19 cs.CV cs.AI 84%

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

基于扩散的多模态3D物体检测在恶劣天气中的修复

Zhijian He, Feifei Liu, Yuwei Li, Zhanpeng Luo, Jintao Cheng, Xieyuanli Chen, Xiaoyu Tang

机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15747 2025-12-19 cs.LG cs.CL cs.CV cs.CY 84%

D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models

D3G:多样化的人口数据生成提高多模态模型中的零样本图像分类准确性

Javon Hickmon

机构 * Javon Hickmon(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 D3G通过生成多样化的人口数据,减少多模态模型中的偏见,提高零样本图像分类的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16270 2025-12-19 cs.CV cs.AI 73%

TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering

TextEditBench: 评估超出渲染的推理-aware文本编辑

Rui Gui, Yang Wan, Haochen Han, Dongxing Mao, Fangming Liu, Min Li, Alex Jinpeng Wang

机构 * Central South University(中南大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 TextEditBench 旨在评估文本编辑中超出渲染的推理能力,通过引入语义期望维度,推动多模态生成中的文本引导编辑技术发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06833 2025-12-19 cs.CV 70%

ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search

ConsistTalk: 可控强度的时序一致说话头生成与扩散噪声搜索

Zhenjie Liu, Jianzhang Lu, Renjie Lu, Cong Liang, Shangfei Wang

机构 * The corresponding author.(通讯作者)

专题命中 多模态生成 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 ConsistTalk通过引入光流引导时间模块、音频到强度模型和扩散噪声初始化策略,实现了可控强度和时序一致的说话头生成,有效减少闪烁并提升音频视频同步质量。

Comments AAAI26 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00878 2025-12-19 cs.CL cs.AI 62%

Less is More: Resource-Efficient Low-Rank Adaptation

少即是多:资源高效的低秩适应

Chunlin Tian, Xuyang Wei, Huanrong Liu, Zhijiang Guo, Li Li

机构 * University of Macau(澳门大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 EffiLoRA通过统一A矩阵和动态B矩阵更新,在多种模型中实现更高效且鲁棒的低秩适应。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16910 2025-12-19 cs.CV cs.LG 57%

SFTok: Bridging the Performance Gap in Discrete Tokenizers

SFTok: 缩小离散分词器在性能上的差距

Qihang Rao, Borui Zhang, Wenzhao Zheng, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 SFTok通过多步骤迭代机制和自引导视觉重建策略,提升图像重建质量,实现高压缩率下的高性能表现。

Comments Under review. Code is available at https://github.com/Neur-IO/SFTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16841 2025-12-19 cs.CV 57%

Radiology Report Generation with Layer-Wise Anatomical Attention

基于层间解剖注意力的放射报告生成

Emmanuel D. Muñiz-De-León, Jorge A. Rosales-de-Golferichs, Ana S. Muñoz-Rodríguez, Alejandro I. Trejo-Castro, Eduardo de Avila-Armenta, Antonio Martínez-Torteya

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于层间解剖注意力的紧凑图像到文本模型,用于生成胸部X光报告的发现部分,通过整合解剖区域信息提升了报告的准确性和连贯性。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16776 2025-12-19 cs.CV 57%

Kling-Omni Technical Report

Kling-Omni 技术报告

Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, Zipeng Feng, Kun Gai, Sainan Guo, Feng Han, Jingbin He, Kang He, Xiao Hu, Xiaohua Hu, Boyuan Jiang, Fangyuan Kong, Hang Li, Jie Li, Qingyu Li, Shen Li, Xiaohan Li, Yan Li, Jiajun Liang, Borui Liao, Yiqiao Liao, Weihong Lin, Quande Liu, Xiaokun Liu, Yilun Liu, Yuliang Liu, Shun Lu, Hangyu Mao, Yunyao Mao, Haodong Ouyang, Wenyu Qin, Wanqi Shi, Xiaoyu Shi, Lianghao Su, Haozhi Sun, Peiqin Sun, Pengfei Wan, Chao Wang, Chenyu Wang, Meng Wang, Qiulin Wang, Runqi Wang, Xintao Wang, Xuebo Wang, Zekun Wang, Min Wei, Tiancheng Wen, Guohao Wu, Xiaoshi Wu, Zhenhua Wu, Da Xie, Yingtong Xiong, Yulong Xu, Sile Yang, Zikang Yang, Weicai Ye, Ziyang Yuan, Shenglong Zhang, Shuaiyu Zhang, Yuanxing Zhang, Yufan Zhang, Wenzheng Zhao, Ruiliang Zhou, Yan Zhou, Guosheng Zhu, Yongjie Zhu

机构 * Kling Team(Kling团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 Kling-Omni 是一种通用生成框架,通过多模态视觉语言输入直接合成高质量视频,整合视频生成、编辑和智能推理任务,实现电影级内容创作。

Comments Kling-Omni Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2512.16485 2025-12-19 cs.CV cs.AI 81%

Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors

脸上微笑,眼中悲伤:通过眼和面部行为的多模态数据集弥合情感差距

Kejun Liu, Yuanyuan Liu, Lin Wei, Chang Tang, Yibing Zhan, Zijing Chen, Zhe Chen

机构 * School of Computer Science, China University of Geosciences (Wuhan)(中国地质大学(武汉)计算机科学学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Computer Science, Wuhan University(武汉大学计算机科学学院) School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算科学、工程与数学科学学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文通过构建包含眼行为和面部行为的多模态数据集EMER,提出EMERT模型以提升情感识别的鲁棒性。

Comments Accepted by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏