arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2407.09033 2024-08-01 cs.CV 57%

Textual Query-Driven Mask Transformer for Domain Generalized Segmentation

Byeonghyun Pak, Byeongju Woo, Sunghwan Kim, Dae-hwan Kim, Hoseong Kim

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21438 2024-08-01 cs.CV 57%

A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain Gap

Lijun Zhang, Wei Suo, Peng Wang, Yanning Zhang

专题命中 视觉定位与Grounding :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06709 2024-08-01 cs.CV 57%

AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One

Mike Ranzinger, Greg Heinrich, Jan Kautz, Pavlo Molchanov

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

Comments CVPR 2024 Version 3: CVPR Camera Ready, reconfigured full paper, table 1 is now more comprehensive Version 2: Added more acknowledgements and updated table 7 with more recent results. Ensured that the link in the abstract to our code is working properly Version 3: Fix broken hyperlinks

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 12490-12500

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03181 2024-07-31 cs.AI cs.CL cs.IR 57%

C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models

Mintong Kang, Nezihe Merve Gürel, Ning Yu, Dawn Song, Bo Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted to ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19849 2024-07-30 cs.CV 57%

Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models

Jihun Yi, Dahuin Jung, Sungroh Yoon

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19118 2024-07-30 cs.AI 57%

Large Language Models as Co-Pilots for Causal Inference in Medical Studies

Ahmed Alaa, Rachael V. Phillips, Emre Kıcıman, Laura B. Balzer, Mark van der Laan, Maya Petersen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18244 2024-07-26 cs.CV 57%

RefMask3D: Language-Guided Transformer for 3D Referring Segmentation

Shuting He, Henghui Ding

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ACM MM 2024, Code: https://github.com/heshuting555/RefMask3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16696 2024-07-24 cs.CV 57%

PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects

Junyi Li, Junfeng Wu, Weizhi Zhao, Song Bai, Xiang Bai

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ECCV2024, homepage: https://provencestar.github.io/PartGLEE-Vision/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16638 2024-07-24 cs.CV 57%

Unveiling and Mitigating Bias in Audio Visual Segmentation

Peiwen Sun, Honggang Zhang, Di Hu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ACM MM 24 (ORAL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16127 2024-07-24 cs.CL cs.AI 57%

Finetuning Generative Large Language Models with Discrimination Instructions for Knowledge Graph Completion

Yang Liu, Xiaobin Tian, Zequn Sun, Wei Hu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted in the 23rd International Semantic Web Conference (ISWC 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14715 2024-07-23 cs.CV cs.CL 57%

FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

Hang Hua, Jing Shi, Kushal Kafle, Simon Jenni, Daoan Zhang, John Collomosse, Scott Cohen, Jiebo Luo

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13642 2024-07-19 cs.CV 57%

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

Xiaoyu Zhu, Hao Zhou, Pengfei Xing, Long Zhao, Hao Xu, Junwei Liang, Alexander Hauptmann, Ting Liu, Andrew Gallagher

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11335 2024-07-19 cs.CV 57%

LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction

Penghui Du, Yu Wang, Yifan Sun, Luting Wang, Yue Liao, Gang Zhang, Errui Ding, Yan Wang, Jingdong Wang, Si Liu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13043 2024-07-19 cs.CV 57%

When Do We Not Need Larger Vision Models?

Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, Trevor Darrell

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments Code: https://github.com/bfshi/scaling_on_scales

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12442 2024-07-18 cs.CV 57%

ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference

Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to ECCV 2024. code available at https://github.com/mc- lan/ClearCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12276 2024-07-18 cs.CV 57%

VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation

Zhen Qu, Xian Tao, Mukesh Prasad, Fei Shen, Zhengtao Zhang, Xinyi Gong, Guiguang Ding

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12016 2024-07-18 cs.CL cs.AI 57%

LLM-based Frameworks for API Argument Filling in Task-Oriented Conversational Systems

Jisoo Mok, Mohammad Kachuee, Shuyang Dai, Shayan Ray, Tara Taghavi, Sungroh Yoon

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11503 2024-07-17 cs.CV 57%

Beyond Mask: Rethinking Guidance Types in Few-shot Segmentation

Shijie Chang, Youwei Pang, Xiaoqi Zhao, Lihe Zhang, Huchuan Lu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11351 2024-07-17 cs.CV 57%

Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities

Xu Zheng, Yuanhuiyi Lyu, Lin Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11015 2024-07-17 cs.CL cs.AI 57%

Does ChatGPT Have a Mind?

Simon Goldstein, Benjamin A. Levinstein

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05231 2024-07-17 cs.CV 57%

PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection

Xiaofan Li, Zhizhong Zhang, Xin Tan, Chengwei Chen, Yanyun Qu, Yuan Xie, Lizhuang Ma

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10873 2024-07-16 cs.NE cs.AI 57%

Understanding the Importance of Evolutionary Search in Automated Heuristic Design with Large Language Models

Rui Zhang, Fei Liu, Xi Lin, Zhenkun Wang, Zhichao Lu, Qingfu Zhang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by the 18th International Conference on Parallel Problem Solving From Nature (PPSN 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05645 2024-07-16 cs.CV 57%

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng, Yujin Sun, Yuenan Hou, Xinge Zhu, Sibei Yang, Yuexin Ma

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09781 2024-07-16 cs.CV 57%

Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding

Ruihuang Li, Zhengqiang Zhang, Chenhang He, Zhiyuan Ma, Vishal M. Patel, Lei Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08934 2024-07-15 cs.LG 57%

Compositional Structures in Neural Embedding and Interaction Decompositions

Matthew Trager, Alessandro Achille, Pramuditha Perera, Luca Zancato, Stefano Soatto

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07427 2024-07-15 cs.CV 57%

Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation

Hao Fang, Peng Wu, Yawei Li, Xinxin Zhang, Xiankai Lu

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08268 2024-07-12 cs.CV 57%

Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation

Tong Shao, Zhuotao Tian, Hang Zhao, Jingyong Su

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ECCV24 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06780 2024-07-10 cs.CV 57%

CoLA: Conditional Dropout and Language-driven Robust Dual-modal Salient Object Detection

Shuang Hao, Chunlin Zhong, He Tang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05610 2024-07-09 cs.CV 57%

Described Spatial-Temporal Video Detection

Wei Ji, Xiangyan Liu, Yingfei Sun, Jiajun Deng, You Qin, Ammar Nuwanna, Mengyao Qiu, Lina Wei, Roger Zimmermann

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04068 2024-07-08 cs.CV 57%

CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-aware Prompting

Qinkai Yu, Jianyang Xie, Anh Nguyen, He Zhao, Jiong Zhang, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV

Comments Accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏