arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2112.07133 2023-05-12 cs.CV 57%

CLIP-Lite: Information Efficient Visual Representation Learning with Language Supervision

Aman Shrivastava, Ramprasaath R. Selvaraju, Nikhil Naik, Vicente Ordonez

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.06227 2023-05-04 cs.LG math.PR math.ST stat.ML stat.TH 57%

Convergence for score-based generative modeling with polynomial complexity

Holden Lee, Jianfeng Lu, Yixin Tan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 43 pages

Journal ref Advances in Neural Information Processing Systems 35 (2022), 22870--22882

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08491 2023-04-18 cs.CV 57%

Delving into Shape-aware Zero-shot Semantic Segmentation

Xinyu Liu, Beiwen Tian, Zhen Wang, Rui Wang, Kehua Sheng, Bo Zhang, Hao Zhao, Guyue Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2023, code: https://github.com/Liuxinyv/SAZS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08408 2023-04-18 cs.CV 57%

OVTrack: Open-Vocabulary Multiple Object Tracking

Siyuan Li, Tobias Fischer, Lei Ke, Henghui Ding, Martin Danelljan, Fisher Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10914 2023-04-12 cs.CV cs.CL 57%

Prophet Attention: Predicting Attention with Future Attention for Image Captioning

Fenglin Liu, Xuancheng Ren, Xian Wu, Wei Fan, Yuexian Zou, Xu Sun

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03669 2023-04-10 cs.CV 57%

DATE: Domain Adaptive Product Seeker for E-commerce

Haoyuan Li, Hao Jiang, Tao Jin, Mengyan Li, Yan Chen, Zhijie Lin, Yang Zhao, Zhou Zhao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments This paper was accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12048 2023-04-07 cs.LG cs.SI stat.ML 57%

Ollivier-Ricci Curvature for Hypergraphs: A Unified Framework

Corinna Coupette, Sebastian Dalleiger, Bastian Rieck

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted at ICLR 2023 (https://openreview.net/forum?id=sPCKNl5qDps)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16891 2023-03-30 cs.CV 57%

Mask-free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask Annotations

Vibashan VS, Ning Yu, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M. Patel, Ran Xu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2023. Project site: https://vibashan.github.io/ovis-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14338 2023-03-28 cs.CV 57%

Turning a CLIP Model into a Scene Text Detector

Wenwen Yu, Yuliang Liu, Wei Hua, Deqiang Jiang, Bo Ren, Xiang Bai

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13431 2023-03-28 cs.CV cs.RO 57%

Instruction-Following Agents with Multimodal Transformer

Hao Liu, Lisa Lee, Kimin Lee, Pieter Abbeel

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments fixed a typo in affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13542 2023-03-27 cs.AI cs.DL math.HO 57%

OntoMath${}^{\mathbf{PRO}}$ 2.0 Ontology: Updates of the Formal Model

Alexander Kirillovich, Olga Nevzorova, Evgeny Lipachev

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11755 2023-03-22 cs.CV 57%

LIMITR: Leveraging Local Information for Medical Image-Text Representation

Gefen Dawidowicz, Elad Hirsch, Ayellet Tal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10962 2023-03-21 cs.RO cs.CV 57%

Neural Implicit Vision-Language Feature Fields

Kenneth Blomqvist, Francesco Milano, Jen Jen Chung, Lionel Ott, Roland Siegwart

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14563 2023-03-20 cs.CV cs.CL 57%

Who are you referring to? Coreference resolution in image narrations

Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01046 2023-03-16 cs.CV 57%

Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos

Daizong Liu, Pan Zhou

专题命中 视觉定位与Grounding :visual reasoning(abstract);分类 cs.CV

Comments Accepted by ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08007 2023-03-15 cs.RO cs.AI 57%

Continuous Risk Measures for Driving Support

Julian Eggert, Tim Puphal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref International Journal of Automotive Engineering (IJAE 2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05646 2023-03-13 cs.CV 57%

Iterative Few-shot Semantic Segmentation from Image Label Text

Haohan Wang, Liang Liu, Wuhao Zhang, Jiangning Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang, Haoqian Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ijcai 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04229 2023-03-09 cs.AI cs.CL 57%

Understanding Natural Language Understanding Systems. A Critical Analysis

Alessandro Lenci

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments to appear in Sistemi Intelligenti

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05499 2023-03-07 cs.CV 57%

CLIP the Gap: A Single Domain Generalization Approach for Object Detection

Vidit Vidit, Martin Engilberge, Mathieu Salzmann

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00632 2023-03-02 cs.CY cs.AI cs.SC 57%

That's All Folks: a KG of Values as Commonsense Social Norms and Behaviors

Stefano De Giorgis, Aldo Gangemi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 16 pages, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14163 2023-03-01 cs.CV 57%

A Language-Guided Benchmark for Weakly Supervised Open Vocabulary Semantic Segmentation

Prashant Pandey, Mustafa Chasmai, Monish Natarajan, Brejesh Lall

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05093 2023-02-21 cs.IR cs.AI cs.MM 57%

Unified Vision-Language Representation Modeling for E-Commerce Same-Style Products Retrieval

Ben Chen, Linbo Jin, Xinxin Wang, Dehong Gao, Wen Jiang, Wei Ning

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

Comments Accepted in The Web Conference (WWW2023) Industry Track. 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02633 2023-02-07 cs.AI 57%

Toward a normative theory of (self-)management by goal-setting

Nishad Singhi, Florian Mohnert, Ben Prystawski, Falk Lieder

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00268 2023-02-02 cs.CV 57%

Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

Kaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao, Qianru Sun

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments accepted by ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11087 2023-01-27 cs.AI 57%

Generalized Planning as Heuristic Search: A new planning search-space that leverages pointers over objects

Javier Segovia-Aguas, Sergio Jiménez, Anders Jonsson

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Under review in the Artificial Intelligence Journal (AIJ)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10283 2023-01-26 cs.CL cs.LG 57%

Audience-Centric Natural Language Generation via Style Infusion

Samraj Moorjani, Adit Krishnan, Hari Sundaram, Ewa Maslowska, Aravind Sankar

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 14 pages, 3 figures, Accepted in Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.08980 2023-01-24 cs.SE cs.AI 57%

Towards Quantification of Assurance for Learning-enabled Components

Erfan Asaadi, Ewen Denney, Ganesh Pai

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 8 pp, 4 figures, Appears in the proceedings of EDCC 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07336 2023-01-19 cs.CV 57%

Class Enhancement Losses with Pseudo Labels for Zero-shot Semantic Segmentation

Son Duy Dao, Hengcan Shi, Dinh Phung, Jianfei Cai

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.08829 2023-01-18 cs.AI cs.RO 57%

Symbol Emergence in Cognitive Developmental Systems: a Survey

Tadahiro Taniguchi, Emre Ugur, Matej Hoffmann, Lorenzo Jamone, Takayuki Nagai, Benjamin Rosman, Toshihiko Matsuka, Naoto Iwahashi, Erhan Oztop, Justus Piater, Florentin Wörgötter

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 23 pages, 6 figures. Submitted to IEEE Transactions on Cognitive and Developmental Systems

Journal ref IEEE Transactions on Cognitive and Developmental Systems, vol. 11, no. 4, pp. 494-516, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11323 2023-01-10 cs.LG stat.ML 57%

A Generalized EigenGame with Extensions to Multiview Representation Learning

James Chapman, Ana Lawry Aguila, Lennie Wells

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏