arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2306.09683 2024-05-24 cs.CV 57%

Scaling Open-Vocabulary Object Detection

Matthias Minderer, Alexey Gritsenko, Neil Houlsby

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12200 2024-05-21 cs.CV 57%

Multi-View Attentive Contextualization for Multi-View 3D Object Detection

Xianpeng Liu, Ce Zheng, Ming Qian, Nan Xue, Chen Chen, Zhebin Zhang, Chen Li, Tianfu Wu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11315 2024-05-21 cs.CV 57%

MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection

Ximiao Zhang, Min Xu, Dehui Qiu, Ruixin Yan, Ning Lang, Xiuzhuang Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 12 pages, 3 figures, 5 tables, early accepted at MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10832 2024-05-20 cs.CV 57%

Open-Vocabulary Spatio-Temporal Action Detection

Tao Wu, Shuqiu Ge, Jie Qin, Gangshan Wu, Limin Wang

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14646 2024-05-17 cs.LG stat.ML 57%

More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory

James B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail Belkin

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Appeared in ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09431 2024-05-16 cs.CV cs.GR 57%

A Survey On Text-to-3D Contents Generation In The Wild

Chenhan Jiang

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments 11 pages, 10 figures, 4 tables. arXiv admin note: text overlap with arXiv:2401.17807 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08593 2024-05-15 cs.CV 57%

Open-Vocabulary Object Detection via Neighboring Region Attention Alignment

Sunyuan Qiang, Xianfei Li, Yanyan Liang, Wenlong Liao, Tao He, Pai Peng

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06586 2024-05-13 cs.CV 57%

Enhancing Weakly Supervised Semantic Segmentation with Multi-modal Foundation Models: An End-to-End Approach

Elham Ravanbakhsh, Cheng Niu, Yongqing Liang, J. Ramanujam, Xin Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04782 2024-05-09 cs.CV 57%

Dual-Image Enhanced CLIP for Zero-Shot Anomaly Detection

Zhaoxiang Zhang, Hanqiu Deng, Jinan Bao, Xingyu Li

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13007 2024-05-08 cs.AI 57%

A Critical Survey on Fairness Benefits of Explainable AI

Luca Deck, Jakob Schoeffer, Maria De-Arteaga, Niklas Kühl

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01233 2024-05-03 q-fin.MF cs.LG q-fin.CP 57%

Mathematics of Differential Machine Learning in Derivative Pricing and Hedging

Pedro Duarte Gomes

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00195 2024-04-30 cs.CV 57%

Raising the Bar of AI-generated Image Detection with CLIP

Davide Cozzolino, Giovanni Poggi, Riccardo Corvi, Matthias Nießner, Luisa Verdoliva

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06209 2024-04-26 cs.CV 57%

Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, Saining Xie

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Project page: https://tsb0601.github.io/mmvp_blog/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12015 2024-04-19 cs.CV 57%

What does CLIP know about peeling a banana?

Claudia Cuttano, Gabriele Rosi, Gabriele Trivigno, Giuseppe Averta

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to MAR Workshop at CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10864 2024-04-18 cs.CV 57%

Vocabulary-free Image Classification and Semantic Segmentation

Alessandro Conti, Enrico Fini, Massimiliano Mancini, Paolo Rota, Yiming Wang, Elisa Ricci

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Under review, 22 pages, 10 figures, code is available at https://github.com/altndrr/vicss. arXiv admin note: text overlap with arXiv:2306.00917

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11281 2024-04-16 cs.LG stat.ME 57%

Towards Characterizing Domain Counterfactuals For Invertible Latent Causal Models

Zeyu Zhou, Ruqi Bai, Sean Kulinski, Murat Kocaoglu, David I. Inouye

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments In ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12132 2024-04-16 cs.LG 57%

Can Public Large Language Models Help Private Cross-device Federated Learning?

Boxin Wang, Yibo Jacky Zhang, Yuan Cao, Bo Li, H. Brendan McMahan, Sewoong Oh, Zheng Xu, Manzil Zaheer

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Published at Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05426 2024-04-12 cs.CV 57%

Test-Time Zero-Shot Temporal Action Localization

Benedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang, Elisa Ricci

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13034 2024-04-11 cs.CL cs.LG 57%

Subspace Representations for Soft Set Operations and Sentence Similarities

Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted at NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17048 2024-04-10 cs.CV 57%

Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions

Zeyu Han, Fangrui Zhu, Qianru Lao, Huaizu Jiang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments CVPR 2024, Code available at https://github.com/Show-han/Zeroshot_REC

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05687 2024-04-09 cs.CV 57%

Retrieval-Augmented Open-Vocabulary Object Detection

Jooyeon Kim, Eulrang Cho, Sehyung Kim, Hyunwoo J. Kim

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted paper at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05937 2024-04-09 cs.CV 57%

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Chengjian Feng, Yujie Zhong, Zequn Jie, Weidi Xie, Lin Ma

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04883 2024-04-09 cs.CV 57%

Mixture of Low-rank Experts for Transferable AI-Generated Image Detection

Zihan Liu, Hanyi Wang, Yaoyu Kang, Shilin Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07633 2024-04-09 cs.CL cs.LG 57%

Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model

Yizhou Zhang, Loc Trinh, Defu Cao, Zijun Cui, Yan Liu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

Comments 9 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04231 2024-04-08 cs.CV 57%

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03650 2024-04-05 cs.CV 57%

OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

Francis Engelmann, Fabian Manhardt, Michael Niemeyer, Keisuke Tateno, Marc Pollefeys, Federico Tombari

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments ICLR 2024, Project page: https://opennerf.github.io

Journal ref ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02132 2024-04-05 cs.CV 57%

ViTamin: Designing Scalable Vision Models in the Vision-Language Era

Jieneng Chen, Qihang Yu, Xiaohui Shen, Alan Yuille, Liang-Chieh Chen

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR 2024; https://github.com/Beckschen/ViTamin

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11782 2024-04-04 cs.CV 57%

Learning Object State Changes in Videos: An Open-World Perspective

Zihui Xue, Kumar Ashutosh, Kristen Grauman

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by CVPR 2024, Project website: https://vision.cs.utexas.edu/projects/VidOSC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01491 2024-04-03 cs.CV 57%

SUGAR: Pre-training 3D Visual Representations for Robotics

Shizhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to CVPR 2024. Project webpage: https://cshizhe.github.io/projects/robot_sugar.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17922 2024-04-03 cs.CV 57%

A Simple Recipe for Language-guided Domain Generalized Segmentation

Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez, Raoul de Charette

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏