arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2507.04769 2025-07-08 cs.CV cs.AI 62%

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Ying Deng, Zhiqiang Yuan, Jiapei Zhang, Jinchao Zhang, Jie Zhou

机构 * WeChat AI, Tencent Inc(腾讯公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09721 2025-07-08 cs.LG cs.CV 62%

Finetuning CLIP to Reason about Pairwise Differences

Dylan Sam, Devin Willmott, Joao D. Semedo, J. Zico Kolter

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03275 2025-07-08 cs.CV cs.LG 62%

ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization

Haosheng Gan, Berk Tinaz, Mohammad Shahab Sepehri, Zalan Fabian, Mahdi Soltanolkotabi

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

Comments An earlier version appeared in the CVPR 2025 Workshop on Generative Models for Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24016 2025-07-01 cs.CL cs.AI cs.CV 62%

EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations

Hyunjong Kim, Sangyeop Kim, Jongheon Jeong, Yeongjae Cho, Sungzoon Cho

机构 * Seoul National University(首尔国立大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23639 2025-07-01 cs.CV cs.AI 62%

Unified Multimodal Understanding via Byte-Pair Visual Encoding

Wanpeng Zhang, Yicheng Feng, Hao Luo, Yijiang Li, Zihao Yue, Sipeng Zheng, Zongqing Lu

机构 * Peking University(北京大学) UC San Diego(加州大学圣地亚哥分校) Renmin University of China(中国人民大学) BeingBeyond

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21144 2025-06-27 cs.LG cs.CV 62%

Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion

Yuguang Zhang, Kuangpu Guo, Zhihe Lu, Yunbo Wang, Jian Liang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Hamad Bin Khalifa University(哈马德·本·卡西姆大学) Central South University(中南大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15724 2025-06-23 cs.LG cs.AI cs.CL 62%

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference

Kunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang, Chengfei Lv, Shengyu Zhang, Fan Wu, Fei Wu

机构 * Zhejiang University(浙江大学) Southeast University(东南大学) Alibaba(阿里巴巴) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03147 2025-06-23 cs.CV cs.AI cs.CL 62%

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Bin Lin, Zongjian Li, Xinhua Cheng, Yuwei Niu, Yang Ye, Xianyi He, Shenghai Yuan, Wangbo Yu, Shaodong Wang, Yunyang Ge, Yatian Pang, Li Yuan

机构 * Peking University(北京大学) Shenzhen Graduate School(深圳研究生院) Peng Cheng Laboratory(鹏城实验室) Rabbitpre AI

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14797 2025-06-19 cs.LG cs.AI 62%

Bound by semanticity: universal laws governing the generalization-identification tradeoff

Marco Nurisso, Jesseba Fernando, Raj Deshpande, Alan Perotti, Raja Marjieh, Steven M. Frankland, Richard L. Lewis, Taylor W. Webb, Declan Campbell, Francesco Vaccarino, Jonathan D. Cohen, Giovanni Petri

机构 * Dipartimento di Scienze Matematiche, Politecnico di Torino(都灵理工大学数学科学系) CENTAI Institute(CENTAI研究院) Network Science Institute, Northeastern University(东北大学网络科学研究所) Institute for Experiential AI, Northeastern University(东北大学体验人工智能研究所) NP Lab, Network Science Institute, Northeastern University London(东北大学伦敦网络科学研究所NP实验室) Department of Psychology, Princeton University(普林斯顿大学心理学系) Program in Cognitive Science, Dartmouth College(达特茅斯学院认知科学项目) Department of Psychology, University of Michigan(密歇根大学心理学系) Microsoft Research(微软研究院) Princeton Neuroscience Institute(普林斯顿神经科学研究所) Department of Physics, Northeastern University(东北大学物理系)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11906 2025-06-18 cs.CV cs.AI 62%

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension

Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments This is the camera-ready version for ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04280 2025-06-11 cs.RO cs.AI cs.LG 62%

Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models

Niccolò Turcato, Matteo Iovino, Aris Synodinos, Alberto Dalla Libera, Ruggero Carli, Pietro Falco

机构 * Department of Information Engineering, University of Padova(帕多瓦大学信息工程系) ABB Corporate Research(ABB企业研究)

专题命中 其他VLM :visual language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07399 2025-06-10 cs.CV cs.AI 62%

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04788 2025-06-06 cs.CL cs.AI cs.LG 62%

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

Jisu An, Junseok Lee, Jeoungeun Lee, Yongseok Son

机构 * Seoul National University(首尔国立大学) University of California San Diego(加州大学圣地亚哥分校) Chung-Ang University(Chung-Ang 大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04129 2025-06-06 eess.IV cs.AI cs.CV 62%

Recent Advances in Medical Image Classification

Loan Dao, Ngoc Quoc Ly

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

Journal ref International Journal of Advanced Computer Science and Applications(ijacsa), 15(7), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15228 2025-06-04 cs.LG cs.CV 62%

Learning from True-False Labels via Multi-modal Prompt Retrieving

Zhongnian Li, Jinghao Xu, Peng Ying, Meng Wei, Xinzheng Xu

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13945 2025-06-03 cs.AI cs.CL cs.LG 62%

CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks

Jie Feng, Jun Zhang, Tianhui Liu, Xin Zhang, Tianjian Ouyang, Junbo Yan, Yuwei Du, Siqi Guo, Yong Li

机构 * Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系、北京理工大学、清华大学) School of Electronic and Information Engineering, Beijing Jiaotong University(电子信息工程学院、北京交通大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院、清华大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG

Comments Accepted by KDD 2025 D&B Track, https://github.com/tsinghua-fib-lab/CityBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14604 2025-06-02 cs.CV cs.AI cs.CL 62%

Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

Sara Sarto, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) IIT-CNR(意大利国家研究委员会IIT)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments IJCAI 2025. Repo GitHub: https://github.com/aimagelab/awesome-captioning-evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22200 2025-05-29 cs.CV cs.AI 62%

Investigating Mechanisms for In-Context Vision Language Binding

Darshana Saravanan, Makarand Tapaswi, Vineet Gandhi

机构 * CVIT, IIIT Hyderabad, India(计算机视觉研究所,印度海得拉巴印度理工学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to MIV at CVPRW 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23907 2025-05-29 cs.CV cs.AI 62%

HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment

Zhichao Liao, Xiaokun Liu, Wenyu Qin, Qingyu Li, Qiulin Wang, Pengfei Wan, Di Zhang, Long Zeng, Pingfa Feng

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21459 2025-05-28 cs.DB cs.AI cs.CV cs.IR cs.MM 62%

LazyVLM: Neuro-Symbolic Approach to Video Analytics

Xiangru Jian, Wei Pang, Zhengyuan Dong, Chao Zhang, M. Tamer Özsu

机构 * University of Waterloo(滑铁卢大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures, Working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18770 2025-05-27 cs.CV cs.LG 62%

Dual-Path Stable Soft Prompt Generation for Domain Generalization

Yuedi Zhang, Shuanghao Bai, Wanqi Zhou, Zhirong Luan, Badong Chen

机构 * Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学) School of Electrical Engineering(电气工程学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00275 2025-05-21 cs.CV cs.AI 62%

Exploring Social Media Image Categorization Using Large Models with Different Adaptation Methods: A Case Study on Cultural Nature's Contributions to People

Rohaifa Khaldi, Domingo Alcaraz-Segura, Ignacio Sánchez-Herrera, Javier Martinez-Lopez, Carlos Javier Navarro, Siham Tabik

机构 * Dept. of Computer Science and Artificial Intelligence, DaSCI, University of Granada(计算机科学与人工智能系,DaSCI,格拉纳达大学) Interuniversity Institute of Earth System Research in Andalusia (IISTA), University of Granada(安达卢西亚地球系统研究中心(IISTA),格拉纳达大学) Dept. of Botany, University of Granada(植物学系,格拉纳达大学) Andalusian Center for Global Change (ENGLOBA), University of Almería(安达卢西亚全球变化中心(ENGLOBA),阿尔梅里亚大学) Dept. of Ecology, University of Granada(生态学系,格拉纳达大学) EDUCA EDTECH Group, Camino de la Torrecilla, 30, 18220, Granada, Spain(EDUCA EDTECH集团,Torrecilla路30号,格拉纳达,西班牙)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03607 2025-05-19 cs.AI cs.CL cs.CV cs.CY cs.HC 62%

Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning

Trilok Padhi, Ugur Kursuncu, Yaman Kumar, Valerie L. Shalin, Lane Peterson Fronczek

机构 * Georgia State University(佐治亚州立大学) Adobe MDSR Wright State University(怀特州立大学) California Polytechnic State University(加州州立大学帕克校区)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at IEEE International Conference on Big Data 2024 (IEEE BigData 2024)

Journal ref IEEE International Conference on Big Data 2024 (IEEE BigData 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07251 2025-05-13 cs.CV cs.AI 62%

Incomplete In-context Learning

Wenqiang Wang, Yangshijie Zhang

机构 * Sun Yat-sen University(中山大学) Lanzhou University(兰州大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03703 2025-05-07 cs.CV cs.LG 62%

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

François Role, Sébastien Meyer, Victor Amblard

机构 * Université Paris-Cité(巴黎-cite大学) Pôle d’Expertise de la Régulation Numérique (PEReN)(数字监管专家中心)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10316 2025-05-06 cs.CV cs.AI 62%

BrushEdit: All-In-One Image Inpainting and Editing

Yaowei Li, Yuxuan Bian, Xuan Ju, Zhaoyang Zhang, Junhao Zhuang, Ying Shan, Yuexian Zou, Qiang Xu

机构 * Peking University(北京大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments WebPage available at https://liyaowei-stu.github.io/project/BrushEdit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18538 2025-04-28 cs.LG cs.AI cs.RO 62%

Generalization Capability for Imitation Learning

Yixiao Wang

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13351 2025-04-21 cs.RO cs.AI cs.HC cs.LG cs.MM 62%

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

Chen Wang, Fei Xia, Wenhao Yu, Tingnan Zhang, Ruohan Zhang, C. Karen Liu, Li Fei-Fei, Jie Tan, Jacky Liang

机构 * Google DeepMind(谷歌DeepMind) Stanford University(斯坦福大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.AI、cs.LG

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14142 2025-04-17 cs.CV cs.AI 62%

Imagery as Inquiry: Exploring A Multimodal Dataset for Conversational Recommendation

Se-eun Yoon, Hyunsik Jeon, Julian McAuley

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10775 2025-04-17 cs.CV cs.AI cs.MA 62%

COMBO: Compositional World Models for Embodied Multi-Agent Cooperation

Hongxin Zhang, Zeyuan Wang, Qiushi Lyu, Zheyuan Zhang, Sunli Chen, Tianmin Shu, Behzad Dariush, Kwonjoon Lee, Yilun Du, Chuang Gan

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

Comments Published at ICLR 2025. 24 pages. The first three authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏