arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-12 至 2026-02-12 共收录 71 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 17 篇

2602.11124 2026-02-12 cs.CV 79%

PhyCritic: Multimodal Critic Models for Physical AI

PhyCritic:面向物理AI的多模态批评模型

Tianyi Xiong, Shihao Wang, Guilin Liu, Yi Dong, Ming Li, Heng Huang, Jan Kautz, Zhiding Yu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 PhyCritic通过双阶段RLVR流程优化,提升物理AI任务中的感知和推理能力,实现对物理任务的高效评估和改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11847 2026-02-12 cs.IR cs.AI cs.CY 79%

A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches

多模态制造安全聊天机器人:知识库设计、基准开发及多种RAG方法的评估

Ryan Singh, Austin Hamilton, Amanda White, Michael Wise, Ibrahim Yousif, Arthur Carvalho, Zhe Shan, Reza Abrisham Baf, Mohammad Mayyas, Lora A. Cavuoto, Fadel M. Megahed

机构 * Farmer School of Business, Miami University(Miami大学农业商学院) Department of Computer Science and Software Engineering, Miami University(Miami大学计算机科学与软件工程系) Department of Engineering Technology, Miami University(Miami大学工程技术系) Department of Mechanical and Manufacturing Engineering, Miami University(Miami大学机械与制造工程系) Department of Industrial and Systems Engineering, University at Buffalo(University at Buffalo工业与系统工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种多模态制造安全聊天机器人,通过RAG方法实现高准确率、低延迟和低成本的安全培训,展示了其在工业5.0环境中的应用价值。

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17667 2026-02-12 cs.AI 79%

PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level

PhysUniBench: 一个面向本科生层面的多模态物理推理基准测试

Lintao Wang, Encheng Su, Jiaqi Liu, Pengze Li, Jiabei Xiao, Wenlong Zhang, Xinnan Dai, Xi Chen, Yuan Meng, Lei Bai, Wanli Ouyang, Shixiang Tang, Aoran Wang, Xinzhu Ma

机构 * The University of Sydney(悉尼大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学) Michigan State University(密歇根州立大学) Tsinghua University(清华大学) Beihang University(北航大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 PhysUniBench是一个面向本科生层面的多模态物理推理基准测试,旨在评估和提升多模态大语言模型在物理问题上的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03220 2026-02-12 eess.SP 78%

Multimodal-Wireless: A Large-Scale Dataset for Sensing and Communication

多模态无线:一种大规模数据集用于传感与通信

Tianhao Mao, Le Liang, Jie Yang, Hao Ye, Shi Jin, Geoffrey Ye Li

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出多模态无线数据集,用于多模态传感与通信研究,包含高分辨率CSI与多种传感器数据,支持通信与协同感知应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16471 2026-02-12 cs.CV 70%

Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos

秩序从混沌:从 glitchy 游戏视频中理解物理世界

Meng Cao, Haoran Tang, Haoze Zhao, Mingfei Han, Ruyang Liu, Qiang Sun, Xiaojun Chang, Ian Reid, Xiaodan Liang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Peking University(北京大学) University of Toronto(多伦多大学) Sun Yat-sen University(孙中山大学)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出 PhysGame 数据集和 GameBench 基准,通过利用游戏视频中的 glitch 异常,提升物理推理能力,实验显示在多个基准上取得显著提升。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15933 2026-02-12 cs.CV 70%

City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs

城市真实环境中的导航:探索从大规模知识中涌现的导航

Dwip Dalal, Utkarsh Mishra, Narendra Ahuja, Nebojsa Jojic

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出稀疏关联视觉导航任务,通过CityNav基准评估MLLMs在真实城市环境中的导航能力,并提出VoP方法提升导航性能。

Comments Accepted at EACL 2026 (ORAL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11144 2026-02-12 cs.LG cs.AI cs.CV 62%

GENIUS: Generative Fluid Intelligence Evaluation Suite

GENIUS:生成性流体智能评估套件

Ruichuan An, Sihan Yang, Ziyu Guo, Wei Dai, Zijun Shen, Haodong Li, Renrui Zhang, Xinyu Wei, Guopeng Li, Wenshan Wu, Wentao Zhang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 GENIUS通过评估生成性流体智能,揭示模型在动态推理和上下文适应上的不足,并提出无需训练的注意力干预策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10675 2026-02-12 cs.CV cs.AI 62%

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning

TwiFF (Think With Future Frames): 一个大规模动态视觉推理数据集

Junhua Liu, Zhangcheng Wang, Zhike Han, Ningli Wang, Guotao Liang, Kun Kuang

机构 * College of Computer Science and Technology, Zhejiang University, Hangzhou, China(浙江大学计算机科学与技术学院) School of Computer and Computing Science, Hangzhou City University, Hangzhou, China(杭州市城市大学计算机与计算科学学院) Henan Academy of Innovations in Medical Science (AIMS), Zhengzhou, China(河南省医学创新研究院) School of Software, Beihang University, Beijing, China(北京航空航天大学软件学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 TwiFF提出一个大规模动态视觉推理数据集及模型,通过整合视频生成与图像理解能力,提升动态场景下的视觉问答性能。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10575 2026-02-12 cs.CV cs.AI cs.CY 62%

MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning

MetaphorStar: 基于端到端视觉强化学习的图像隐喻理解与推理

Chenhao Zhang, Yazhe Niu, Hongsheng Li

机构 * Shanghai AI Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong MMLab(香港中文大学MMLab)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MetaphorStar通过端到端视觉强化学习框架提升图像隐喻理解与推理能力,显著优于现有模型。

Comments 14 pages, 4 figures, 11 tables; Code: https://github.com/MING-ZCH/MetaphorStar, Model & Dataset: https://huggingface.co/collections/MING-ZCH/metaphorstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.11496 2026-02-12 eess.IV cs.CV 57%

MITI: SLAM Benchmark for Laparoscopic Surgery

MITI:腹腔手术的SLAM基准测试

Regine Hartwig, Daniel Ostler, Jean-Claude Rosenthal, Hubertus Feußner, Dirk Wilhelm, Dirk Wollherr

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 MITI基准测试提供了一套用于评估微创手术中SLAM算法性能的高质量数据集,包含多模态传感器信息和校准数据,旨在推动视觉-惯性算法的发展。

Comments This submission is withdrawn because it is a duplicate of "Constrained Visual-Inertial Localization With Application And Benchmark in Laparoscopic Surgery" (arXiv:2202.11075). The withdrawn version contains less complete information. Readers are directed to the full version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10518 2026-02-12 cs.CV 57%

MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps

MapVerse:一个用于在多样化真实世界地图上进行地理问答的基准测试

Sharat Bhat, Harshita Khandelwal, Tushar Kataria, Vivek Gupta

机构 * University of Southern California(南加州大学) University of California Los Angeles(加州大学洛杉矶分校) University of Utah(犹他大学) Arizona State University(亚利桑那州立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 MapVerse是一个基于真实世界地图的大规模基准测试,用于评估地理问答任务中的多模态推理能力,揭示了当前VLMs在复杂空间推理任务上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19899 2026-02-12 cs.CV 57%

VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering

VeriSciQA:一个用于科学视觉问答的自动验证数据集

Yuyi Li, Daoyuan Chen, Zhen Wang, Yutong Lu, Yaliang Li

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 VeriSciQA通过跨模态验证框架生成并验证高质量科学视觉问答数据集,显著提升开源模型在科学图表问答任务上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10529 2026-02-12 cs.CY 50%

Drawing Your Programs: Exploring the Applications of Visual-Prompting with GenAI for Teaching and Assessment

绘制你的程序:探索视觉提示与生成式AI在教学与评估中的应用

David H. Smith, S. Moonwara A. Monisha, Annapurna Vadaparty, Leo Porter, Daniel Zingaro

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文探讨利用视觉提示与生成式AI在编程教学与评估中的应用,通过学生构建的分解图提示GPT-4.1生成代码,展示多模态提示在编程教育中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10407 2026-02-12 cs.HC cs.LG 50%

Towards Affordable, Non-Invasive Real-Time Hypoglycemia Detection Using Wearable Sensor Signals

迈向低成本、非侵入式实时低血糖检测的可穿戴传感器信号

Lawrence Obiuwevwi, Krzysztof J. Rechowicz, Vikas Ashok, Sampath Jayarathna

机构 * Department of Computer Science, Old Dominion University, Norfolk, VA, USA(计算机科学系,旧 Dominion 大学) Virginia Digital Maritime Center (VDMC), Old Dominion University, Norfolk, VA, USA(弗吉尼亚数字水路中心(VDMC),旧 Dominion 大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本研究提出了一种多模态深度学习方法,利用可穿戴传感器信号实现低成本、非侵入式实时低血糖检测,通过融合皮肤电反应和心率信号提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2602.10814 2026-02-12 cs.AI 79%

See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch

见、计划、快照:评估Scratch中的多模态GUI代理

Xingyi Zhang, Yulei Ye, Kaifeng Huang, Wenhao Li, Xiangfeng Wang

机构 * Tongji University(同济大学) East China Normal University(华东师范大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出ScratchWorld基准,用于评估多模态GUI代理在Scratch中的程序构建能力,揭示了推理与执行之间的显著差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11057 2026-02-12 cs.LG 78%

Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models

分割、调和、然后征服:利用多模态语言模型解决多商品流问题

Xinyu Yuan, Yan Qiao, Zonghui Wang, Wenzhi Chen

机构 * Zhejiang University(浙江大学) Hefei University of Technology(合肥工业大学) Co-corresponding authors(共同通讯作者)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 Pram利用多模态语言模型解决多商品流问题,通过分解和调和子问题实现高效优化,性能接近线性规划求解器且运行时间显著降低。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10648 2026-02-12 cs.HC 71%

Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied Knowledge

生成性肌肉刺激:通过将多模态AI与具身知识相结合,为用户提供物理帮助

Yun Ho, Romain Nith, Peili Jiang, Steven He, Bruno Felalaga, Shan-Yuan Teng, Rhea Seeralan, Pedro Lopes

专题命中 多模态Agent :multimodal(title)

AI总结 本研究提出通过结合多模态AI与具身知识生成肌肉刺激指令,实现更通用的物理辅助系统。

Comments 22 pages, 29 figures

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 13 篇

2509.14671 2026-02-12 cs.CL cs.AI cs.LG 86%

TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding

TableDART: 表格理解的动态自适应多模态路由

Xiaobo Xing, Wei Yuan, Tong Chen, Quoc Viet Hung Nguyen, Xiangliang Zhang, Hongzhi Yin

机构 * The University of Queensland, Australia(昆士兰大学) Griffith University, Australia(格里菲斯大学) University of Notre Dame, USA(诺丁汉大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);cross-modal(abstract)

AI总结 TableDART通过动态选择文本、图像或融合视角,提升表格理解的准确性和效率,避免昂贵的多模态模型微调。

Comments Accepted to ICLR 2026. 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10138 2026-02-12 cs.CV cs.AI cs.CL cs.LG 85%

Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement

多模态信息融合用于图表理解:MLLMs的综述——演变、局限与认知增强

Zhihang Yi, Jian Zhao, Jiancheng Lv, Tao Wang

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Engineering Research Center of Machine Learning(机器学习工程研究中心) China Telecom Institute of AI(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了多模态大语言模型在图表理解中的应用,分析了其演变、局限及未来发展方向,旨在推动更稳健的系统发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22463 2026-02-12 cs.MM 83%

Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation

正交解缠与投影特征对齐用于对话中多模态情绪识别

Xinyi Che, Wenbo Wang, Jian Guan, Qijun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出OD-PFA框架,通过正交解缠与投影特征对齐技术提升对话中多模态情绪识别性能。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02172 2026-02-12 cs.CV cs.AI cs.MM 82%

GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting

GaussianCross: 通过高斯点撒技术实现跨模态自监督3D表示学习

Lei Yao, Yi Wang, Yi Zhang, Moyun Liu, Lap-Pui Chau

机构 * Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 GaussianCross通过高斯点撒技术实现跨模态自监督3D表示学习,提升3D点云的表示质量和泛化能力。

Comments 14 pages, 8 figures, accepted by MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10161 2026-02-12 cs.CR cs.AI cs.CL 79%

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐

Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出OmniSteer方法,通过提取黄金拒绝向量和轻量级适配器提升多模态模型的安全性与通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12768 2026-02-12 cs.CV cs.AI cs.LG 73%

CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence

CoRe3D:协作推理作为3D智能的基础

Tianjiao Yu, Xinzhuo Li, Yifan Shen, Yuanzhe Liu, Ismini Lourentzou

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 CoRe3D通过协作推理框架实现3D内容生成,结合语义和空间推理提升3D输出的一致性与描述一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 70%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09651 2026-02-12 cs.CV cs.AI cs.LG 62%

Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era

地理空间表示学习:从深度学习到大语言模型时代的一次综述

Xixuan Hao, Yutian Jiang, Xingchen Zou, Jiabo Liu, Yifang Yin, Song Gao, Flora Salim, Tianrui Li, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The University of New South Wales(新南威尔士大学) Southwest Jiaotong University(西南交通大学) Institute for Infocomm Research (I$^2$R), A*STAR(信息通信研究院(I$^2$R),A*STAR)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了从深度学习到大语言模型时代的地理空间表示学习,探讨了其方法、应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10845 2026-02-12 cs.AI cs.LG 57%

SynergyKGC: Reconciling Topological Heterogeneity in Knowledge Graph Completion via Topology-Aware Synergy

SynergyKGC: 通过拓扑感知协同解决知识图谱补全中的拓扑异质性

Xuecheng Zou, Yu Tang, Bingbing Wang

机构 * School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) School of Mathematical Sciences, Soochow University(数学科学学院,苏州大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 SynergyKGC通过拓扑感知协同解决知识图谱补全中的拓扑异质性问题,提升KGC命中率。

Comments 10 pages, 5 tables, 7 figures. This work introduces the Active Synergy mechanism and Identity Anchoring for Knowledge Graph Completion. Code: https://github.com/XuechengZou-2001/SynergyKGC-main

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10740 2026-02-12 cs.CL 57%

Reinforced Curriculum Pre-Alignment for Domain-Adaptive VLMs

强化课程预对齐用于领域自适应视觉-语言模型

Yuming Yan, Shuo Yang, Kai Tang, Sihong Chen, Yang Zhang, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu, Edith C. H. Ngai

机构 * Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出RCPA方法,通过课程意识的渐进调节机制,在领域自适应中平衡领域知识获取与通用能力保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10494 2026-02-12 cs.CL 57%

Canvas-of-Thought: Grounding Reasoning via Mutable Structured States

Canvas-of-Thought:通过可变的结构化状态进行推理

Lingzhuang Sun, Yuxia Zhu, Ruitong Liu, Hao Liang, Zheng Sun, Caijun Jia, Honghao He, Yuchen Wu, Siyuan Li, Jingxuan Wei, Xiangxiang Zhang, Bihui Yu, Wentao Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) New York University(纽约大学) Westlake University(西湖大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 Canvas-CoT通过引入HTML Canvas实现可变结构化状态,提升多模态推理效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10053 2026-02-12 cs.CV 57%

DiCo: Disentangled Concept Representation for Text-to-image Person Re-identification

DiCo: 用于文本到图像人物重识别的解耦概念表示

Giyeol Kim, Chanho Eom

机构 * organization= Department of Imaging Science, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea organization= Department of Metaverse Convergence, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 DiCo通过解耦概念表示方法,提升文本到图像人物重识别的跨模态对齐和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15147 2026-02-12 cs.CV 57%

From Pixels to Images: A Structural Survey of Deep Learning Paradigms in Remote Sensing Image Semantic Segmentation

从像素到图像:深度学习在遥感图像语义分割中的结构调查

Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang

机构 * College of Science and Engineering and Centre for AI and Data Science Innovation, James Cook University(科学与工程学院和人工智能与数据科学创新中心,詹姆斯库克大学) Department of Forest and Wildlife Ecology, University of Wisconsin-Madison(森林与野生动物生态学系,威斯康星大学麦迪逊分校) School of Computing, Engineering and Mathematical Sciences, La Trobe University(计算、工程与数学科学学院,拉特罗布大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文系统回顾了深度学习在遥感图像语义分割中的结构演变,从像素到图像的层次化方法,涵盖多种技术并提供可复现的代码库。

Comments 34 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏