arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2508.03562 2025-08-06 cs.CV cs.CL 57%

Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching

Muzhaffar Hazman, Susan McKeever, Josephine Griffith

机构 * School of Computer Science University of Galway(计算机科学学院 Galway大学) School of Computer Science Technological University Dublin(计算机科学学院 技术大学都柏林)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted for publication at IEEE International Conference on Image Processing Theory, Tools and Applications (IPTA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03535 2025-08-06 cs.CV 57%

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu, Wenshuo Chen, Yutao Yue

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00991 2025-08-06 cs.CV 57%

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

Xiaorong Zhu, Ziheng Jia, Jiarui Wang, Xiangyu Zhao, Haodong Duan, Xiongkuo Min, Jia Wang, Zicheng Zhang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06690 2025-08-05 cs.SD cs.AI cs.MM 57%

Benchmarking Sub-Genre Classification For Mainstage Dance Music

Hongzhi Shu, Xinglin Li, Hongyu Jiang, Minghao Fu, Xinyu Li

机构 * Johns Hopkins University(约翰霍普金斯大学) Southeast University(东南大学) National University of Defense Technology(国防科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01074 2025-08-05 cs.CV cs.CR 57%

Evading Data Provenance in Deep Neural Networks

Hongyu Zhu, Sichu Liang, Wenwen Wang, Zhuomeng Zhang, Fangqi Li, Shi-Lin Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Southeast University(东南大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04981 2025-08-05 cs.CV 57%

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting

Xiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei, Nan Dong, Ming-Hsuan Yang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) Chongqing Changan Automobile Co., Ltd(重庆长安汽车有限公司) University of California, Merced(加州大学默塞德分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 Hightlight (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00391 2025-08-04 cs.CV eess.AS 57%

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22075 2025-07-31 cs.LG 57%

Prototype-Guided Pseudo-Labeling with Neighborhood-Aware Consistency for Unsupervised Adaptation

Eman Ali, Chetan Arora, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Indian Institute of Technology Delhi(印度德里理工学院) Alexandria University(亚历山大大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21745 2025-07-29 cs.CV 57%

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu, Tengfei Wang, Gordon Wetzstein, Dahua Lin, Ziwei Liu

机构 * Fudan University(复旦大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Stanford University(斯坦福大学) The Chinese University of Hong Kong(香港中文大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments Page: https://zyh482.github.io/3DGen-Bench/ ; Code: https://github.com/3DTopia/3DGen-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15680 2025-07-24 cs.CV 57%

Visual-Language Model Knowledge Distillation Method for Image Quality Assessment

Yongkang Hou, Jiarun Song

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22139 2025-07-23 cs.CV 57%

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Shaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo, Jian Luan

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10908 2025-07-23 cs.CV 57%

Do large language vision models understand 3D shapes?

Sagi Eppel

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13684 2025-07-23 cs.CV 57%

FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models

Minghan Li, Chenxi Xie, Yichen Wu, Lei Zhang, Mengyu Wang

机构 * Harvard AI and Robotics Lab, Harvard University(哈佛人工智能与机器人实验室,哈佛大学) Broad Institute(博德研究所) Hong Kong Polytechnic University(香港理工大学) School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) City University of Hong Kong(香港城市大学) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(自然与人工智能研究学院,哈佛大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 24 pages, 14 figures, 16 tables

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15569 2025-07-22 cs.CV 57%

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团) Peking University(北京大学) Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15025 2025-07-22 cs.SE cs.AI 57%

Survey of GenAI for Automotive Software Development: From Requirements to Executable Code

Nenad Petrovic, Vahid Zolfaghari, Andre Schamschurko, Sven Kirchner, Fengjunjie Pan, Chengdng Wu, Nils Purschke, Aleksei Velsh, Krzysztof Lebioda, Yinglei Song, Yi Zhang, Lukasz Mazur, Alois Knoll

机构 * Chair of Robotics, Artificial Intelligence and Real-Time Systems(机器人、人工智能与实时系统教授席)

专题命中 其他VLM :vision language model(abstract);分类 cs.AI

Comments Conference paper accepted for GACLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00302 2025-07-22 cs.LG cond-mat.mtrl-sci 57%

Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework

Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban

机构 * Computer Engineering, Texas A\&M University, College Station, TX 77843, USA College of Science Engineering, Hamad Bin Khalifa University, Doha, Qatar Dept. of Electrical \& Computer Engineering, Texas A\&M University at Qatar, Doha, Qatar Orthotics, Ankara University, Ankara, Turkey

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments Presented at ICML 2025 Workshop on DataWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23765 2025-07-18 cs.CV 57%

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) China University of Geosciences(中国地质大学) Nanyang Technological University(南洋理工大学) BAAI(百度人工智能研究院) Stanford University(斯坦福大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09961 2025-07-15 cs.LG 57%

Text-Driven Causal Representation Learning for Source-Free Domain Generalization

Lihua Zhou, Mao Ye, Nianxin Li, Shuaifeng Li, Jinlin Wu, Xiatian Zhu, Lei Deng, Hongbin Liu, Jiebo Luo, Zhen Lei

机构 * Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究所,中国科学院,香港,中国) School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学) Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究所,CVSSP, Surrey大学) School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学) University of Rochester(罗切斯特大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10563 2025-07-15 cs.CV 57%

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Jiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang, Yubo Wang, Yuansheng Ni, Wang Zhu, Ziyan Jiang, Bohan Lyu, Dongfu Jiang, Xuan He, Yuan Liu, Hexiang Hu, Xiang Yue, Wenhu Chen

机构 * Core Contributors(核心贡献者) Tiger-AI-Lab(虎鲸人工智能实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR 2025 camera-ready version. Project page: https://tiger-ai-lab.github.io/MEGA-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07105 2025-07-10 cs.CV eess.IV 57%

4KAgent: Agentic Any Image to 4K Super-Resolution

Yushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang, Renjie Li, Jian Wang, Yide Zhang, Gengchen Mai, Lihong V. Wang, James Zou, Xiaoyu Wang, Ming-Hsuan Yang, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) Stanford University(斯坦福大学) Snap Inc.(Snap公司) CU Boulder(科罗拉多大学博尔德分校) UT Austin(得克萨斯大学奥斯汀分校) California Institute of Technology(加州理工学院) Topaz Labs(Topaz实验室) UC Merced(加州大学默塞德分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Project page: https://4kagent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06973 2025-07-10 cs.CV 57%

Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

Qiyuan Dai, Sibei Yang

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Sun Yat-sen University(孙中山大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06523 2025-07-10 cs.CV cs.CL cs.GR 57%

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui, Xinya Du

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05344 2025-07-08 cs.CV 57%

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Jiahui Wang, Zuyan Liu, Yongming Rao, Jiwen Lu

机构 * Tsinghua University(清华大学) Tencent Hunyuan X(腾讯混元实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03997 2025-07-04 cs.CV 57%

CAD-Editor: A Locate-then-Infill Framework with Automated Training Data Synthesis for Text-Based CAD Editing

Yu Yuan, Shizhao Sun, Qi Liu, Jiang Bian

机构 * University of Science(科学大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06285 2025-07-04 cs.CV 57%

DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation

Yueming Lyu, Tianwei Lin, Fu Li, Dongliang He, Jing Dong, Tieniu Tan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CRIPAC, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所CRIPAC) VIS, Baidu Inc.(百度公司VIS)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Code is available at https://github.com/Yueming6568/DeltaEdit

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00907 2025-07-02 cs.CR cs.AI 57%

The Age of Sensorial Zero Trust: Why We Can No Longer Trust Our Senses

Fabio Correa Xavier

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00891 2025-07-02 cs.CL cs.AI 57%

MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes

Yuheng Wang, Xianhe Tang, Pufeng Huang

机构 * Wuhan University(武汉大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00586 2025-07-02 cs.CV 57%

Context-Aware Academic Emotion Dataset and Benchmark

Luming Zhao, Jingwen Xuan, Jiamin Lou, Yonghui Yu, Wenwu Yang

机构 * Zhejiang Gongshang University(浙江工商大学) Zhejiang Yuexiu University(浙江越秀大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05769 2025-07-02 cs.CV 57%

Exploring Text-Guided Single Image Editing for Remote Sensing Images

Fangzhou Han, Lingyu Si, Zhizhuo Jiang, Hongwei Dong, Lamei Zhang, Yu Liu, Hao Chen, Bo Du

机构 * Department of Information Engineering, Harbin Institute of Technology(信息工程系,哈尔滨工业大学) National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(空间信息集成国家重点实验室,中国科学院软件研究所) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Hubei Luojia Laboratory, National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(湖北珞珈实验室,国家多媒体软件工程技术研究中心,武汉大学计算机学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 17 pages, 18 figures, Accepted by IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21476 2025-06-27 cs.CV 57%

Global and Local Entailment Learning for Natural World Imagery

Srikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal, Nathan Jacobs

机构 * Washington University in St. Louis(圣路易斯华盛顿大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏