arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2401.11401 2024-01-23 cs.CV 70%

LLMRA: Multi-modal Large Language Model based Restoration Assistant

Xiaoyu Jin, Yuan Shi, Bin Xia, Wenming Yang

专题命中 其他VLM :vision language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00698 2023-10-03 cs.CV cs.HC 70%

Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips

Reshma Ramaprasad

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted at CLVL: 5th Workshop On Closing The Loop Between Vision And Language (ICCV 2023 Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04588 2026-08-25 cs.CL 版本更新 67%

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

VCIFBench:评估视频理解中的复杂指令遵循能力

Huangchen Xu, Yuan Wu, Yi Chang

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University(知识驱动人机智能工程研究中心,吉林大学) International Center of Future Science, Jilin University(未来科学国际中心,吉林大学)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract_cn)

AI总结 提出VCIFBench基准,通过混合验证流水线评估多模态大模型在视频理解中遵循内容、格式、风格和结构约束的复杂指令能力,实验表明联合约束满足仍具挑战,DPO训练可提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19207 2026-08-21 cs.CL 新提交 67%

Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

合规性、能力与冲突:基于系统消息的多模态大语言模型基准测试

Juan Yeo, Geewook Kim

机构 * NAVER Cloud(NAVER云) KAIST AI(韩国科学技术院人工智能)

专题命中 其他VLM :multimodal large language model(abstract,abstract_cn)

AI总结 该研究构建基准VSysBench测试多模态大语言模型在系统消息下的合规性,发现施加系统消息会降低任务准确率,开放权重模型易受用户冲突影响,视觉约束最难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27938 2026-07-31 cs.HC 新提交 67%

VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs

VizPilot:基于多模态大语言模型的SVG复合可视化自动引导系统

Nishaanthini Gnanavel, Yong Wang

专题命中 其他VLM :multimodal large language model(abstract,abstract_cn)

AI总结 VizPilot是基于多模态大语言模型的SVG复合可视化自动引导工具,通过双模块实现自动生成交互式引导,经评估可降低创作工作量并减轻用户认知负荷,提升复合可视化可用性。

Comments Accepted by IEEE VIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05121 2026-06-30 cs.CL cs.SD eess.AS 67%

The NTNU System at the S&I Challenge 2025 SLA Open Track

NTNU系统在S&I挑战2025 SLA开放赛道中的表现

Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

机构 * National Taiwan Normal University(国立台湾师范大学) Department of Computer Science and Information Engineering(计算机科学与信息工程系) Institute of AI Interdisciplinary Applied Technology(人工智能跨学科应用技术研究所)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

AI总结 本文提出整合W2V与Phi-4多模态大语言模型的系统,通过分数融合策略提升SLA评估性能,在Speak & Improve Challenge 2025中取得第二名,RMSE为0.375。

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23611 2026-06-23 cs.CV cs.AI cs.LG 新提交 67%

Data Selection Through Iterative Self-Filtering for Vision-Language Settings

面向视觉-语言场景的迭代自过滤数据选择

Andrei Liviu Nicolicioiu, Sarvjeet Singh Ghotra, Morgane M. Moss, Aaron Courville

机构 * Mila, Université de Montréal(蒙特利尔大学Mila实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出一种自举式方法,通过迭代训练CLIP模型并自选择改进的数据混合,无需额外数据或预训练模型即可提升下游性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16848 2026-05-19 cs.CV cs.AI cs.CL cs.LG 67%

Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction

基于模式的思考:通过模式诱导突破视觉规划中的感知瓶颈

Yichang Jian, Boyuan Xiao, Zhenyuan Huang, Yifei Peng, Yao-Xiang Ding

机构 * State Key Lab of CAD& CG(CAD与CG国家重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出通过模式诱导的方法,利用模式推理和模式诱导策略,使视觉语言模型在视觉规划任务中实现更高效和准确的感知与推理,解决传统模型在复杂输入下的感知瓶颈问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10012 2026-05-12 cs.HC cs.CR 67%

Sketch-based Access Control: A Multimodal Interface for Translating User Preferences into Intent-Aligned Policies

基于草图的访问控制:一种多模态接口,用于将用户偏好转化为意图对齐的策略

Kyzyl Monteiro, Sauvik Das

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract_cn)

AI总结 本文提出SBAC系统,结合草图和多模态大语言模型,帮助用户逐步完善访问策略,发现潜在问题并验证策略行为。

Comments 27 pages including appendix; 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02200 2026-04-22 cs.NE cs.AI cs.CV cs.LG 67%

Learning Evolution via Optimization Knowledge Adaptation

通过优化知识适应学习进化

Chao Wang, Lingling Li, Licheng Jiao, Jiaxuan Zhao, Fang Liu, Shuyuan Yang

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education(教育部智能感知与图像理解重点实验室) International Research Center for Intelligent Perception and Computation(智能感知与计算国际研究中心) Xidian University(西安电子科技大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出OKAEM模型,通过注意力机制参数化进化算子,实现优化知识的预训练和自适应优化,提升进化算法的知识转移与在线适应能力。

Comments This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07833 2026-04-13 cs.CV cs.AI cs.LG 67%

Relational Visual Similarity

关系视觉相似性

Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee, Yuheng Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of California, Los Angeles(加利福尼亚大学洛杉矶分校) Adobe Research(Adobe研究院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出通过关系相似性而非视觉属性来衡量图像相似性,构建了首个基于关系逻辑的图像表示空间,揭示了现有视觉模型在捕捉关系相似性方面的不足。

Comments CVPR 2026 camera-ready; Project page, data, and code: https://thaoshibe.github.io/relsim

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26107 2026-03-30 cs.HC 67%

One Is Not Enough: How People Use Multiple AI Models in Everyday Life

一个不够:人们如何在日常生活中使用多个AI模型

Seunghwa Pyo, Donggun Lee, Jungwoo Rhee, Soobin Park, Youn-kyung Lim

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

AI总结 研究探讨了人们如何在日常生活中协调多个多模态大语言模型,揭示用户构建模型层级和切换策略以优化任务效率与输出可信度。

Comments Accepted as a poster at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16786 2026-03-06 cs.LG cs.AI cs.CV 67%

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

重新审视多模态KV缓存压缩:一种基于频域的异常KV感知方法

Yaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu, Maosen Zhao, Jia Hao, Tao Chen

机构 * College of Future Information Technology, Fudan University(未来信息科技学院,复旦大学) Shanghai Innovation Institute(上海创新研究院) The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhangjiang Laboratory(张江实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出FlashCache,一种基于频域的异常KV感知KV缓存压缩框架,通过保留关键KV对提升解码效率并降低内存使用。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07666 2026-01-01 cs.LG cs.AI cs.CL cs.CV 67%

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

在大语言模型、多模态大语言模型及更广泛的领域中进行模型融合:方法、理论、应用与机遇

Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao

机构 * Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区) Northeastern University China(东北大学) Shenzhen Campus of Sun Yat-sen University China(中山大学深圳校区) Nanyang Technological University Singapore(南洋理工大学) Northeastern University(东北大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Nanyang Technological University(南洋理工大学) Institute for Clarity in Documentation Dublin Ohio USA(文档清晰研究所) Inria Paris-Rocquencourt Rocquencourt France(巴黎-罗quentourt研究所) Rajiv Gandhi University Doimukh Arunachal Pradesh India(拉贾·甘地大学) Tsinghua University Haidian Qu Beijing Shi China(清华大学) Palmer Research Laboratories San Antonio Texas USA(帕勒研究中心) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究中心)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文综述了模型融合的方法、理论、应用及未来方向,提出新的分类方法并探讨其在多个机器学习领域的应用及挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20781 2025-12-25 cs.IR 67%

Soft Filtering: Guiding Zero-shot Composed Image Retrieval with Prescriptive and Proscriptive Constraints

软过滤:通过规劝性与禁止性约束指导零样本复合图像检索

Youjin Jung, Seongwoo Cho, Hyun-seok Min, Sungchul Choi

专题命中 其他VLM :vision-language model(abstract);multimodal large language model(abstract)

AI总结 本文提出SoFT方法,通过规劝性和禁止性约束提升零样本复合图像检索的准确性与鲁棒性。

Comments Accepted to AAAI 2026 Workshop on New Frontiers in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19933 2025-12-24 cs.CL 67%

PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation

PRISM: 一种基于个性的多智能体框架用于社交媒体模拟

Zhixiang Lu, Xueyuan Deng, Yiran Liu, Yulong Li, Qiang Yan, Imran Razzak, Jionglong Su

机构 * University of Liverpool(利物浦大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University College London(伦敦大学学院) Xi'an Jiaotong-Liverpool University(西安交通大学-利物浦大学) Chinese Academy of Sciences(中国科学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

AI总结 PRISM通过结合连续情绪演变与基于个性的决策过程,提供了一种更准确模拟社交媒体中个性驱动意见极化的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26536 2025-11-26 cs.CL cs.AI cs.CV cs.LG cs.RO 67%

OceanGym: A Benchmark Environment for Underwater Embodied Agents

OceanGym: 一个用于水下具身智能体的基准环境

Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen

专题命中 其他VLM :MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 OceanGym通过多模态大语言模型构建首个水下具身智能体基准,揭示水下环境感知与规划的挑战,推动水下自主系统发展。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19518 2025-11-26 cs.CV cs.AI cs.IT cs.LG math.IT 67%

Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning

迈向高效的VLMs:通过自适应结构压缩的信息论驱动压缩

Zhaoqi Xu, Yingying Zhang, Jian Li, Jianwei Guo, Qiannan Zhu, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Zhongtai Securities Institute for Financial Studies, Shandong University(山东大学中泰证券金融研究学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出InfoPrune框架,通过信息论驱动的自适应结构压缩方法,在保持性能的同时显著提升视觉语言模型的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13883 2025-10-17 q-bio.NC cs.MA 67%

Large Language Model Agents Enable Autonomous Design and Image Analysis of Microwell Microfluidics

Dinh-Nguyen Nguyen, Sadia Shakil, Raymond Kai-Yu Tong, Ngoc-Duy Dinh

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21956 2025-09-30 cs.CV cs.AI cs.CL cs.LG 67%

Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation

Mengdan Zhu, Senhao Cheng, Guangji Bai, Yifei Zhang, Liang Zhao

机构 * Emory University(埃默里大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16995 2025-09-23 cs.DC 67%

MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

Zheming Yang, Qi Guo, Yunqing Hu, Chang Zhao, Chang Zhang, Jian Zhao, Wen Ji

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06850 2025-09-16 cs.NE 67%

Visual Evolutionary Optimization on Graph-Structured Combinatorial Problems with MLLMs: A Case Study of Influence Maximization

Jie Zhao, Kang Hao Cheong

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13180 2025-07-25 cs.CV cs.AI cs.LG 67%

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras, Tushar Nagarajan, Muhammad Maaz, Yale Song, Tengyu Ma, Shuming Hu, Suyog Jain, Miguel Martin, Huiyu Wang, Hanoona Rasheed, Peize Sun, Po-Yao Huang, Daniel Bolya, Nikhila Ravi, Shashank Jain, Tammy Stark, Shane Moon, Babak Damavandi, Vivian Lee, Andrew Westbury, Salman Khan, Philipp Krähenbühl, Piotr Dollár, Lorenzo Torresani, Kristen Grauman, Christoph Feichtenhofer

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11798 2025-07-22 cs.CR 67%

BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model

Weilin Lin, Nanjun Zhou, Yanyun Wang, Jianze Li, Hui Xiong, Li Liu

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12378 2025-07-17 cs.IR cs.CL 67%

Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker

Rachna Saxena, Abhijeet Kumar, Suresh Shanmugam

专题命中 其他VLM :visual language model(abstract);MLLM(abstract)

Comments Presented at NLP@IR workshop at SIGIR conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08260 2025-06-25 cs.SE 67%

FixDrive: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation

Yang Sun, Christopher M. Poskitt, Kun Wang, Jun Sun

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by the 47th IEEE/ACM International Conference on Software Engineering (ICSE 2025)

Journal ref Proc. ICSE'25, pages 1921-1933. IEEE, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18095 2025-06-24 cs.CV cs.AI cs.LG 67%

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen, Ke Ji, Xidong Wang, Yunjin Yang, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08023 2025-06-11 q-bio.BM cs.AI cs.CE cs.CV cs.LG 67%

Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Qifeng Wu, Zhengzhe Liu, Han Zhu, Yizhou Zhao, Daisuke Kihara, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 4 pages for body, 3 pages for appendix, 11 figures. Accepted to CVPR 2025 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities(MMFM-BIOMED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07484 2025-06-10 cs.CV cs.AI cs.LG 67%

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization

Dasol Hong, Wooju Lee, Hyun Myung

机构 * Urban Robotics Lab, School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Republic of Korea(乌尔班机器人实验室,电气工程学院,韩国科学技术院,大韩民国)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 8 pages, 5 figures; accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13487 2025-05-23 cs.CL cs.AI cs.CV cs.LG 67%

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

Chen-An Li, Tzu-Han Lin, Yun-Nung Chen, Hung-yi Lee

机构 * National Taiwan University(台湾大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏