arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2512.02713 2025-12-03 cs.AI 57%

Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs

利用本体对齐的知识图谱训练数据归因于图像生成

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

机构 * National Centre for Scientific Research ``Demokritos''(国家科学研究中心「德莫克里特」) University of Glasgow(格拉斯哥大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出利用本体对齐的知识图谱方法,通过多模态大语言模型提取图像中的结构化三元组,以追踪生成模型中训练数据的影响,从而提升透明度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02517 2025-12-03 cs.CV 57%

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

SkyMoE:一种用于增强遥感解释的视觉-语言基础模型

Jiaqi Liu, Ronghao Fu, Lang Sun, Haoran Liu, Xiao Yang, Weipeng Zhang, Xu Na, Zhuoran Duan, Bo Yang

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 SkyMoE通过Mixture-of-Experts架构提升遥感多模态多任务处理能力,实现对不同粒度任务的高效适应与优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06263 2025-12-02 cs.CV 57%

OmniSVG: A Unified Scalable Vector Graphics Generation Model

OmniSVG: 一种统一的可扩展矢量图形生成模型

Yiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng, Fukun Yin, Jiaxu Zhang, Liao Wang, Gang Yu, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) StepFun Project(StepFun项目)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 OmniSVG通过统一框架和预训练视觉-语言模型,实现高效多模态SVG生成,提升复杂结构的表达能力,并引入大规模数据集推动SVG合成发展。

Comments 20 pages; Project Page: https://omnisvg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00557 2025-12-02 cs.CV 57%

NeuroVolve: Evolving Visual Stimuli toward Programmable Neural Objectives

NeuroVolve:通过可编程神经目标演化视觉刺激

Haomiao Chen, Keith W Jamison, Mert R. Sabuncu, Amy Kuceyeski

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技) Weill Cornell Medicine(韦尔医学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 NeuroVolve通过可编程神经目标生成视觉刺激,揭示大脑区域间的协同与对抗性调节关系,实现脑引导的图像编辑与首选刺激生成的统一。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21389 2025-11-27 cs.IR cs.AI 57%

FITRep: Attention-Guided Item Representation via MLLMs

FITRep: 通过大语言模型实现的注意力引导的项目表示

Guoxiao Zhang, Ao Li, Tan Qu, Qianlong Xie, Xingxing Wang

机构 * Meituan(美团)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 FITRep通过引入注意力引导的白盒表示框架,利用多模态大语言模型实现细粒度项目去重,提升了广告点击率和每千次展示成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19200 2025-11-26 cs.CV 57%

Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?

现代视觉模型能否理解物体与相似物之间的差异?

Itay Cohen, Ethan Fetaya, Amir Rosenfeld

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文研究了现代视觉模型能否区分真实物体与相似物,通过构建RoLA数据集并改进CLIP模型的嵌入空间方向,提升跨模态检索和描述生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17136 2025-11-24 cs.SD cs.AI 57%

Device-Guided Music Transfer

设备引导的音乐转移

Manh Pham Hung, Changshuo Hu, Ting Dang, Dong Ma

机构 * Singapore Management University(新加坡管理大学) University of Melbourne(墨尔本大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 DeMT通过提取扬声器频率响应曲线生成设备嵌入,实现跨设备的音乐风格迁移与鲁棒适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17534 2025-11-21 cs.CV cs.CL cs.MM 57%

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

协同强化学习用于统一多模态理解和生成

Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出CoRL框架,通过协同强化学习提升多模态大语言模型在生成与理解任务上的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23584 2025-11-21 cs.CV 57%

VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement

VividFace: 高质量和高效的一步扩散用于视频面部增强

Shulian Zhang, Yong Guo, Long Peng, Ziyang Wang, Ye Chen, Wenbo Li, Xiao Zhang, Yulun Zhang, Jian Chen

机构 * South China University of Technology(华南理工大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

AI总结 VividFace通过单步扩散框架和联合训练策略,高效提升视频面部增强的高质量与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13326 2025-11-20 stat.AP cs.AI 57%

TacEleven: generative tactic discovery for football open play

Siyao Zhao, Hao Ma, Zhiqiang Pu, Jingjing Huang, Yi Pan, Shijie Wang, Zhi Ming

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(交叉科学学院,中国科学院大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shanghai AI Laboratory(上海人工智能实验室) Association de la Jeunesse Auxerroise(亚眠青年协会)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13993 2025-11-19 cs.CV 57%

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments NeurIPS 2025, Project webpage: https://vision.cs.utexas.edu/projects/CrossTrainer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12644 2025-11-18 cs.CV 57%

Use as Many Surrogates as You Want: Selective Ensemble Attack to Unleash Transferability without Sacrificing Resource Efficiency

Bo Yang, Hengwei Zhang, Jindong Wang, Yuchen Ren, Chenhao Lin, Chao Shen, Zhengyu Zhao

机构 * Information Engineering University(信息工程大学) Xi’an Jiaotong University(西安交通大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04743 2025-11-18 cs.CV 57%

SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs

Shuhan Xu, Siyuan Liang, Hongling Zheng, Aishan Liu, Xinbiao Wang, Yong Luo, Fu Lin, Leszek Rutkowski, Dacheng Tao

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12040 2025-11-18 cs.CV 57%

SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images

Xinyuan Hu, Changyue Shi, Chuxiao Yang, Minghao Chen, Jiajun Ding, Tao Wei, Chen Wei, Zhou Yu, Min Tan

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments AAAI2026-Oral. Project Page: https://xinyuanhu66.github.io/SRSplat/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10202 2025-11-17 cs.CL cs.CV 57%

Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval

Shubhashis Roy Dipta, Francis Ferraro

机构 * Department of Computer Science and Electrical Engineering University of Maryland Baltimore County(计算机科学与电气工程系马里兰大学巴尔的摩县)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted in IJCNLP-AACL 2025 (also presented in MAGMAR 2025 at ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08977 2025-11-13 cs.CV 57%

Efficient and Effective In-context Demonstration Selection with Coreset

Zihua Wang, Jiarui Wang, Haiyang Xu, Ming Yan, Fei Huang, Xu Yang, Xiu-Shen Wei, Siya Mi, Yu Zhang

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments This paper is accepted by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01579 2025-11-12 cs.CV 57%

Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning

Lingfeng He, De Cheng, Di Xu, Huaijie Wang, Nannan Wang

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments AAAI-2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18108 2025-11-12 cs.CV 57%

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Corning Inc(康宁公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Journal ref CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06297 2025-11-11 cs.HC cs.AI 57%

Decomate: Leveraging Generative Models for Co-Creative SVG Animation

Jihyeon Park, Jiyoon Myung, Seone Shin, Jungki Son, Joohyung Han

机构 * MODULABS

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments Accepted at the 1st Workshop on Generative and Protective AI for Content Creation (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06195 2025-11-11 cs.HC cs.AI 57%

AI as intermediary in modern-day ritual: An immersive, interactive production of the roller disco musical Xanadu at UCLA

Mira Winick, Naisha Agarwal, Chiheb Boussema, Ingrid Lee, Camilo Vargas, Jeff Burke

机构 * UCLA(加州大学洛杉矶分校)

专题命中 其他VLM :vision language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03137 2025-11-06 cs.AI 57%

Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks

Shipeng Cen, Ying Tan

机构 * School of Intelligence Science Technology, Institute for Artificial Intellignce, Peking University, Beijing, China Technology, Institute for Artificial Intellignce, National Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09013 2025-11-05 cs.CV 57%

Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal

Rongxin Liao, Feng Li, Yanyan Wei, Zenglin Shi, Le Zhang, Huihui Bai, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01678 2025-11-04 cs.CV 57%

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback

Ropeway Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing, Jinwang Wang, Rui Zhao, Yan Xing, Weihua Chen, Fan Wang

机构 * Zhejiang University(浙江大学) DAMO Academy Alibaba Group(阿里云达摩院) Hupan Lab(虎扑实验室) National University of Singapore(新加坡国立大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01593 2025-11-04 cs.CV 57%

Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation

Yizhu Chen, Chen Ju, Zhicheng Wang, Shuai Xiao, Xu Chen, Jinsong Lan, Xiaoyong Zhu, Ying Chen

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Peking University(北京大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21738 2025-11-04 cs.CY cs.CV 57%

From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China

Weihuan Deng, Yaofu Huang, Luan Chen, Xun Li, Yu Gu, Yao Yao

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12709 2025-11-04 cs.IR cs.CV 57%

SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model

Lin Lin, Jiefeng Long, Zhihe Wan, Yuchi Wang, Dingkang Yang, Shuang Yang, Yueyang Yao, Xu Chen, Zirui Guo, Shengqiang Li, Weiran Li, Hanyu Li, Yaling Mou, Yan Qiu, Haiyang Yu, Xiao Liang, Hongsheng Li, Chao Feng

机构 * ByteDance Douyin SAIL Team, CUHK MMLab(字节跳动抖音SAIL团队,CUHK MMLab)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26824 2025-11-03 cs.DL cs.AI cs.IR 57%

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

Magdalena Lederbauer, Siddharth Betala, Xiyao Li, Ayush Jain, Amine Sehaba, Georgia Channing, Grégoire Germain, Anamaria Leonescu, Faris Flaifil, Alfonso Amayuelas, Alexandre Nozadze, Stefan P. Schmid, Mohd Zaki, Sudheesh Kumar Ethirajan, Elton Pan, Mathilde Franckel, Alexandre Duval, N. M. Anoop Krishnan, Samuel P. Gleason

专题命中 其他VLM :vision language model(abstract);分类 cs.AI

Comments 29 pages, 13 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17313 2025-10-28 cs.LG 57%

Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations

Tal Barami, Nimrod Berman, Ilan Naiman, Amos H. Hason, Rotem Ezra, Omri Azencot

机构 * Faculty of Computer and Information Science(计算机与信息科学学院) Ben-Gurion University of the Negev(贝加尔-内盖夫大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments Accepted for the Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21121 2025-10-27 cs.RO cs.AI 57%

Generalizable Hierarchical Skill Learning via Object-Centric Representation

Haibo Zhao, Yu Qi, Boce Hu, Yizhe Zhu, Ziyan Chen, Heng Tian, Xupeng Zhu, Owen Howell, Haojie Huang, Robin Walters, Dian Wang, Robert Platt

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15745 2025-10-27 eess.IV cs.LG 57%

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏