arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6917 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6917 篇

2603.00156 2026-03-03 cs.CV 57%

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP: 用于鲁棒医学图像分割的双向和一致语言-图像处理

Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mustaqeem Khan

机构 * Department of Computer Engineering, University of Kurdistan, Iran(伊朗库尔德大学计算机工程系) Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(伊拉克埃尔比尔库尔德大学人工智能与创新中心) College of Information Technology, United Arab Emirates University, UAE(阿联酋大学信息科技学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 BiCLIP通过双向多模态融合和一致性目标提升医学图像分割的鲁棒性,有效应对标注稀少和临床伪影挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03059 2026-03-02 cs.CV 57%

CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning

CLAP:通过曲率采样和原型学习实现融合3D感知的无监督3D表示学习

Runjian Chen, Hang Zhang, Avinash Ravichandran, Hyoungseob Park, Wenqi Shao, Alex Wong, Ping Luo

机构 * The University of Hong Kong(香港大学) Cruise Yale University(耶鲁大学) Shanghai AI Laboratory(上海人工智能实验室) HKU Shanghai Intelligent Computing Research Center(香港大学上海智能计算研究中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 CLAP通过曲率采样和原型学习,在无监督3D表示学习中实现图像与点云的联合预训练,提升融合3D感知的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22944 2026-02-27 cs.MM 57%

MViR: Multi-View Visual-Semantic Representation for Fake News Detection

MViR:多视图视觉-语义表示用于虚假新闻检测

Haochen Liang, Xinqi Su, Jun Wang, Chaomeng Chen, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.MM

AI总结 MViR通过多视图视觉-语义表示方法提升虚假新闻检测的准确性,结合图像多视角特征与文本信息进行融合分析。

Comments Accepted by ICASSP'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22368 2026-02-27 cs.SE cs.AI 57%

EyeLayer: Integrating Human Attention Patterns into LLM-Based Code Summarization

EyeLayer:将人类注意力模式整合到基于LLM的代码摘要中

Jiahao Zhang, Yifan Zhang, Kevin Leach, Yu Huang

机构 * Vanderbilt University(范德比尔特大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 EyeLayer通过整合人类眼动模式提升LLM代码摘要效果,实现13.17%的BLEU-4提升。

Comments Accepted at the 34th IEEE/ACM International Conference on Program Comprehension (ICPC 2026), April 12-13, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21716 2026-02-26 cs.CV 57%

TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image Detection

TranX-Adapter: 在MLLMs中弥合artifact与语义以实现鲁棒的AI生成图像检测

Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo

机构 * Wuhan University(武汉大学) Tencent YouTu Lab(腾讯YouTu实验室) Singapore Management University(新加坡管理学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 TranX-Adapter通过引入任务感知的最优传输融合和X-Fusion机制,在MLLMs中提升AI生成图像检测的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00795 2026-02-25 cs.CV 57%

DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

DVLA-RL:基于强化学习门控的双层视觉-语言对齐用于少样本学习

Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Zhibin Wu, Yilong Yin

机构 * Software School, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 DVLA-RL通过双层语义构建和强化学习门控注意力,实现少样本学习中视觉与语言的双层次对齐,提升泛化能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20531 2026-02-25 cs.CV 57%

A Lightweight Vision-Language Fusion Framework for Predicting App Ratings from User Interfaces and Metadata

一种轻量级的视觉-语言融合框架,用于从用户界面和元数据预测应用评分

Azrin Sultana, Firoz Ahmed

机构 * Department of Computer Science, American International University–Bangladesh(计算机科学系,美国国际大学-孟加拉)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种轻量级视觉-语言融合框架,通过结合UI和语义信息预测应用评分,实现了高精度的预测效果。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20479 2026-02-25 cs.CV 57%

Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation

路径解耦双曲流匹配用于少样本适应

Lin Li, Ziqi Jiang, Gefan Ye, Zhenqi He, Jiahui Li, Jun Xiao, Kwang-Ting Cheng, Long Chen

机构 * The Hong Kong University of Science(香港科学与技术大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出路径解耦双曲流匹配方法,通过双曲几何解决少样本适应中的路径缠绕问题,实现更高效的特征传输与对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19870 2026-02-24 cs.CV 57%

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

ApET:基于近似误差的令牌压缩用于高效的视觉语言模型

Qiankun Ma, Ziyao Zhang, Haofei Wang, Jie Chen, Zhen Song, Hairong Zheng

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Peng Cheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Harbin Institute of Technology(哈尔滨工业大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 ApET通过近似误差指导的令牌压缩,在不使用注意力机制的情况下高效压缩视觉语言模型的令牌预算,提升推理效率。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19615 2026-02-24 cs.CV 57%

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

清晰可见,自信推理:用于视觉语言模型盲点的即插即用修复方法

Xin Hu, Haomiao Ni, Yunbei Zhang, Jihun Hamm, Zechen Li, Zhengming Ding

机构 * Department of Computer Science, Tulane University(路易斯安那大学计算机科学系) Department of Computer Science, University of Memphis(密苏里大学计算机科学系)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种无需微调的即插即用模块,通过细化视觉标记和丰富文本提示,提升VLM对罕见物体的推理能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12047 2026-02-24 cs.CV 57%

PSGait: Gait Recognition using Parsing Skeleton

PSGait: 基于解析骨架的步态识别

Hangrui Xu, Zhengxian Wu, Chuanrui Zhang, Zhuohong Chen, Zhifang Liu, Peng Jiao, Haoqian Wang

机构 * The Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院) School of Computer Science and Information Engineering, Hefei University of Technology, China(合肥工业大学计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 PSGait通过解析骨架与轮廓融合的方法提升步态识别的准确性和泛化能力,实现15.7%的精度提升。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17536 2026-02-20 physics.acc-ph cs.AI 57%

Toward a Fully Autonomous, AI-Native Particle Accelerator

迈向完全自主的、原生AI粒子加速器

Chris Tennant

机构 * Jefferson Laboratory(杰弗逊实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出通过AI共同设计构建完全自主的粒子加速器,旨在实现自主运行和最大化性能。

Comments 14 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17478 2026-02-20 cs.CV 57%

QuPAINT: Physics-Aware Instruction Tuning Approach to Quantum Material Discovery

QuPAINT:一种面向量子材料发现的物理感知指令微调方法

Xuan-Bac Nguyen, Hoang-Quan Nguyen, Sankalp Pandey, Tim Faltermeier, Nicholas Borys, Hugh Churchill, Khoa Luu

机构 * CVIU Lab, University of Arkansas, USA(Arkansas大学计算机视觉实验室) University of Utah, USA(犹他大学) Department of Physics, University of Arkansas, USA(Arkansas大学物理系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 QuPAINT通过物理感知指令微调方法,提升量子材料发现的多模态表征能力,建立标准化评估基准。

Comments Project page: https://uark-cviu.github.io/projects/qupaint/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16681 2026-02-19 cs.CV 57%

VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detection

VETime: 基于视觉增强的零样本时间序列异常检测

Yingyuan Yang, Tian Lan, Yifei Gao, Yimeng Lu, Wenjun He, Meng Wang, Chenghao Liu, Chen Zhang

机构 * Department of Industrial Engineering, Tsinghua University, Beijing, China(清华大学工业工程系) Lab, Huawei Technologies Ltd, Beijing, China(华为技术有限公司2012实验室) Datadog AI Research(Datadog人工智能研究)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 VETime通过融合视觉与时间模态,提出零样本时间序列异常检测框架,实现高精度定位与低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15857 2026-02-19 cs.CL 57%

Multi-source Heterogeneous Public Opinion Analysis via Collaborative Reasoning and Adaptive Fusion: A Systematically Integrated Approach

通过协同推理与自适应融合进行多源异构公共意见分析:一种系统整合方法

Yi Liu

机构 * Yi Liu School of Software Xi’an Jiaotong University(刘毅 软件学院 西安交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出CRAF框架,通过协同推理与自适应融合整合传统方法与LLMs,提升多源异构公共意见分析的跨平台适应性与性能

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14492 2026-02-18 cs.CL cs.IR 57%

Query as Anchor: Scenario-Adaptive User Representation via Large Language Model

查询作为锚点:通过大语言模型实现场景自适应的用户表示

Jiahao Yuan, Yike Xu, Jinyong Wen, Baokun Wang, Ziyi Gao, Xiaotong Lin, Yun Liu, Xing Fu, Yu Cheng, Yongchao Liu, Weiqiang Wang, Zhongle Xie

机构 * Ant Group(蚂蚁集团) East China Normal University(东华大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

AI总结 通过Query-as-Anchor框架,利用大语言模型实现动态查询感知的用户表示,提升工业级用户建模的场景适应性和性能表现。

Comments 15 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06738 2026-02-17 cs.CL 57%

AWM: Accurate Weight-Matrix Fingerprint for Large Language Models

AWM: 用于大型语言模型的准确权重矩阵指纹

Boyi Zeng, Lin Chen, Ziwei He, Xinbing Wang, Zhouhan Lin

机构 * LUMIA Lab(LUMIA实验室) School of Artificial Intelligence(人工智能学院) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

AI总结 AWM通过基于权重矩阵的无训练指纹方法,利用线性分配问题和无偏中心核对齐相似性,实现对大型语言模型训练来源的可靠识别。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17822 2026-02-16 cs.CV 57%

Easy-Poly: An Easy Polyhedral Framework For 3D Multi-Object Tracking

Easy-Poly: 一种易于使用的多目标跟踪三维多目标跟踪框架

Peng Zhang, Xin Li, Xin Lin, Liang He

机构 * East China Normal University(东华师范大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 Easy-Poly提出了一种基于过滤的三维多目标跟踪框架,通过四个创新方法提升小目标检测和跟踪精度,实现实时高性能表现。

Comments 8 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09843 2026-02-13 cs.CV 57%

Kelix Technical Report

Kelix技术报告

Boyang Ding, Chenglong Chu, Dunju Zang, Han Li, Jiangxia Cao, Kun Gai, Muhao Wei, Ruiming Tang, Shiyao Wang, Siyang Mao, Xinchen Luo, Yahui Liu, Zhixin Ling, Zhuoran Yang, Ziming Li, Chengru Song, Guorui Zhou, Guowang Zhang, Hao Peng, Hao Wang, Jiaxin Deng, Jin Ouyang, Jinghao Zhang, Lejian Ren, Qianqian Wang, Qigen Hu, Tao Wang, Xingmei Wang, Yiping Yang, Zixing Zhang, Ziqi Wang

机构 * Kuaishou Technology(快手科技)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 Kelix是一种完全离散的自回归统一模型,旨在缩小离散和连续视觉表示之间的理解差距。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11672 2026-02-13 cs.CV 57%

U-Net with Hadamard Transform and DCT Latent Spaces for Next-day Wildfire Spread Prediction

具有Hadamard变换和DCT潜在空间的U-Net用于次日野火蔓延预测

Yingyi Luo, Shuaiang Rong, Adam Watts, Ahmet Enis Cetin

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) US Forest Service(美国森林服务局) Pacific Wildland Fire Science Laboratory(太平洋野火科学实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出TD-FusionUNet模型,通过Hadamard和DCT变换提升野火预测精度,以更少参数实现更高效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10845 2026-02-12 cs.AI cs.LG 57%

SynergyKGC: Reconciling Topological Heterogeneity in Knowledge Graph Completion via Topology-Aware Synergy

SynergyKGC: 通过拓扑感知协同解决知识图谱补全中的拓扑异质性

Xuecheng Zou, Yu Tang, Bingbing Wang

机构 * School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) School of Mathematical Sciences, Soochow University(数学科学学院,苏州大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 SynergyKGC通过拓扑感知协同解决知识图谱补全中的拓扑异质性问题,提升KGC命中率。

Comments 10 pages, 5 tables, 7 figures. This work introduces the Active Synergy mechanism and Identity Anchoring for Knowledge Graph Completion. Code: https://github.com/XuechengZou-2001/SynergyKGC-main

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10740 2026-02-12 cs.CL 57%

Reinforced Curriculum Pre-Alignment for Domain-Adaptive VLMs

强化课程预对齐用于领域自适应视觉-语言模型

Yuming Yan, Shuo Yang, Kai Tang, Sihong Chen, Yang Zhang, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu, Edith C. H. Ngai

机构 * Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出RCPA方法,通过课程意识的渐进调节机制,在领域自适应中平衡领域知识获取与通用能力保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10494 2026-02-12 cs.CL 57%

Canvas-of-Thought: Grounding Reasoning via Mutable Structured States

Canvas-of-Thought:通过可变的结构化状态进行推理

Lingzhuang Sun, Yuxia Zhu, Ruitong Liu, Hao Liang, Zheng Sun, Caijun Jia, Honghao He, Yuchen Wu, Siyuan Li, Jingxuan Wei, Xiangxiang Zhang, Bihui Yu, Wentao Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) New York University(纽约大学) Westlake University(西湖大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 Canvas-CoT通过引入HTML Canvas实现可变结构化状态,提升多模态推理效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10053 2026-02-12 cs.CV 57%

DiCo: Disentangled Concept Representation for Text-to-image Person Re-identification

DiCo: 用于文本到图像人物重识别的解耦概念表示

Giyeol Kim, Chanho Eom

机构 * organization= Department of Imaging Science, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea organization= Department of Metaverse Convergence, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 DiCo通过解耦概念表示方法,提升文本到图像人物重识别的跨模态对齐和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15147 2026-02-12 cs.CV 57%

From Pixels to Images: A Structural Survey of Deep Learning Paradigms in Remote Sensing Image Semantic Segmentation

从像素到图像:深度学习在遥感图像语义分割中的结构调查

Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang

机构 * College of Science and Engineering and Centre for AI and Data Science Innovation, James Cook University(科学与工程学院和人工智能与数据科学创新中心,詹姆斯库克大学) Department of Forest and Wildlife Ecology, University of Wisconsin-Madison(森林与野生动物生态学系,威斯康星大学麦迪逊分校) School of Computing, Engineering and Mathematical Sciences, La Trobe University(计算、工程与数学科学学院,拉特罗布大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文系统回顾了深度学习在遥感图像语义分割中的结构演变,从像素到图像的层次化方法,涵盖多种技术并提供可复现的代码库。

Comments 34 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09066 2026-02-11 cs.LG cs.AI 57%

Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning

谱解耦与增强:一种双域对比框架用于表示学习

Jinjin Guo, Yexin Li, Zhichao Huang, Jun Fang, Zhiyuan Liu, Chao Liu, Pengzhang Liu, Qixia Jiang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 SDE通过双域对比损失和谱增强策略,提升多模态表示学习的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14410 2026-02-11 eess.AS 57%

TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation

TTA:跨语言语音表示的转录、翻译与对齐

Wei Liu, Jiahong Li, Yiwen Shao, Dong Yu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 eess.AS

AI总结 TTA通过大规模训练生成跨语言语音表示,提升语音识别与翻译任务的性能,优于Whisper模型。

Comments Accepted by ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08208 2026-02-10 cs.CL cs.HC 57%

LLMs and people both learn to form conventions -- just not with each other

大语言模型和人类都学会形成惯例——但并不是彼此之间

Cameron R. Jones, Agnese Lombardi, Kyle Mahowald, Benjamin K. Bergen

机构 * Department of Psychology, Stony Brook University(心理学系,石溪大学) Department of Cognitive Science, University of California San Diego(认知科学系,加州圣地亚哥大学) Department of Philology, Literature, and Linguistics, University of Pisa(philology、文学与语言学系,比萨大学) Department of Linguistics, University of Texas at Austin(语言学系,德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 研究发现人类和AI在同类型对话中能形成惯例,但人机对话效果较差,表明对话协调需要共同的解释偏见。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07540 2026-02-10 cs.CV cs.LG 57%

LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing

基于LLM的诊断证据对齐用于有限配对情况下的医学视觉-语言预训练

Huimin Yan, Liang Bai, Xian Yang, Long Chen

机构 * Institute of Intelligent Information Processing, Shanxi University(山西大学智能信息处理研究所) Alliance Manchester Business School, The University of Manchester(曼彻斯特大学阿利安斯商学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出LLM引导的诊断证据对齐方法,旨在解决有限配对数据下医学视觉-语言预训练的诊断表示学习问题,通过提取关键诊断证据提升跨模态对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01212 2026-02-10 cs.CV 57%

TSJNet: A Multi-modality Target and Semantic Awareness Joint-driven Image Fusion Network

TSJNet: 一种多模态目标和语义意识联合驱动的图像融合网络

Yuchan Jie, Yushen Xu, Xiaosong Li, Huafeng Li, Haishu Tan, Feiping Nie

机构 * School of Physics and Optoelectronic Engineering, Foshan University(物理与光电工程学院,佛山大学) School of Information Engineering and Automation, Kunming University of Science and Technology(信息工程与自动化学院,昆明理工大学) School of Artificial Intelligence, Optics and Electronics (i0PEN), School of Computer Science, Northwestern Polytechnical University(人工智能、光学与电子学院(i0PEN),计算机科学学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 TSJNet通过联合驱动的多模态图像融合网络提升目标检测与语义分割的性能,实现7.97%和10.88%的精度提升。

详情

展开后加载摘要…

URL PDF HTML 收藏