arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2203.15442 2022-03-30 cs.CV cs.MM 79%

Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding

Jiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang, Xuwu Wang, Ji Zhang, Liang He, Xin Lin

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14940 2022-03-29 cs.CV 79%

Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, Guoqi Li

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13049 2022-03-29 cs.CV 79%

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, Xin Eric Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09905 2022-03-21 cs.CV 79%

Learning Affordance Grounding from Exocentric Images

Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao, Dacheng Tao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06822 2022-03-15 cs.CV cs.CL cs.RO 79%

Grounding Commands for Autonomous Vehicles via Layer Fusion with Region-specific Dynamic Layer Attention

Hou Pong Chan, Mingxi Guo, Cheng-Zhong Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04222 2022-03-15 cs.CV cs.MM 79%

Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs

Kaifeng Gao, Long Chen, Yulei Niu, Jian Shao, Jun Xiao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2022. Code is available at https://github.com/Dawn-LX/VidSGG-BIG. We also won the 1st place of Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021, with a simplified version of our model.(The code for object tracklets generation is available at https://github.com/Dawn-LX/VidVRD-tracklets)

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04281 2022-03-15 cs.CV 79%

Visual Grounding with Transformers

Ye Du, Zehua Fu, Qingjie Liu, Yunhong Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 7 pagrs, 3 figures. Accepted by ICME'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05243 2022-03-11 cs.CV cs.CL cs.MM 79%

A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach

Xiaohan Lan, Yitian Yuan, Xin Wang, Long Chen, Zhi Wang, Lin Ma, Wenwu Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02966 2022-03-08 cs.CV cs.CL 79%

Exploring Optical-Flow-Guided Motion and Detection-Based Appearance for Temporal Sentence Grounding

Daizong Liu, Xiang Fang, Wei Hu, Pan Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2201.00457

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13959 2022-03-01 cs.IR cs.CL cs.LG 79%

Semi-Structured Query Grounding for Document-Oriented Databases with Deep Retrieval and Its Application to Receipt and POI Matching

Geewook Kim, Wonseok Hwang, Minjoon Seo, Seunghyun Park

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments To appear in AAAI-22 Workshop on Knowledge Discovery from Unstructured Data in Financial Services

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08541 2022-01-17 cs.CV 79%

TransVG: End-to-End Visual Grounding with Transformers

Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, Houqiang Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This paper has been accepted by ICCV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02848 2022-01-14 cs.CV cs.IR 79%

Learning Sample Importance for Cross-Scenario Video Temporal Grounding

Peijun Bao, Yadong Mu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.00457 2022-01-04 cs.CV 79%

Exploring Motion and Appearance Information for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou, Yang Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by AAAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.00454 2022-01-04 cs.CV 79%

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng, Zichuan Xu, Pan Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by AAAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.13031 2021-12-28 cs.CV cs.RO 79%

Grounding Linguistic Commands to Navigable Regions

Nivedita Rufus, Kanishk Jain, Unni Krishnan R Nair, Vineet Gandhi, K Madhava Krishna

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Journal ref 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021, pp. 8593-8600

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10138 2021-12-21 cs.AI 79%

Classical Planning as QBF without Grounding (extended version)

Irfansha Shaik, Jaco van de Pol

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04872 2021-12-16 cs.CV cs.MM 79%

Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding

Zhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li, Gangshan Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2022 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.11475 2021-12-07 cs.CV 79%

Self-supervised Learning for Semi-supervised Temporal Language Grounding

Fan Luo, Shaoxiang Chen, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00475 2021-12-02 cs.CV cs.CL 79%

Weakly-Supervised Video Object Grounding via Causal Intervention

Wei Wang, Junyu Gao, Changsheng Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.05717 2021-12-02 cs.CV cs.CL 79%

Relation-aware Video Reading Comprehension for Temporal Language Grounding

Jialin Gao, Xin Sun, Mengmeng Xu, Xi Zhou, Bernard Ghanem

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by EMNLP-21

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07180 2021-11-16 cs.CL cs.LG 79%

Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning

Yizhen Zhang, Minkyu Choi, Kuan Han, Zhongming Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments 10 pages, 7 figures, 1 appendix, to be published in Neurips 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04321 2021-11-09 cs.CV cs.CL 79%

Towards Debiasing Temporal Sentence Grounding in Video

Hao Zhang, Aixin Sun, Wei Jing, Joey Tianyi Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 13 pages, 6 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03438 2021-11-08 cs.RO cs.CL cs.CV 79%

LanguageRefer: Spatial-Language Model for 3D Visual Grounding

Junha Roh, Karthik Desingh, Ali Farhadi, Dieter Fox

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05624 2021-11-01 cs.CV 79%

End-to-end Multi-modal Video Temporal Grounding

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted in NeurIPS 2021. Project page at https://github.com/wenz116/DRFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10540 2021-09-23 cs.CL cs.AI 79%

Awakening Latent Grounding from Pretrained Language Models for Semantic Parsing

Qian Liu, Dejian Yang, Jiahui Zhang, Jiaqi Guo, Bin Zhou, Jian-Guang Lou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted by ACL 2021 Findings. The first three authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11450 2021-09-23 cs.CV 79%

SAT: 2D Semantics Assisted Training for 3D Visual Grounding

Zhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo Luo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2021 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.09028 2021-09-23 cs.CV 79%

A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric

Yitian Yuan, Xiaohan Lan, Xin Wang, Long Chen, Zhi Wang, Wenwu Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08634 2021-09-20 cs.CL cs.AI 79%

Grounding Natural Language Instructions: Can Large Language Models Capture Spatial Information?

Julia Rozanova, Deborah Ferreira, Krishna Dubba, Weiwei Cheng, Dell Zhang, Andre Freitas

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments *Equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08478 2021-09-20 cs.CL cs.CV cs.MM 79%

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li, Jie Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ACL Fingdings 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06400 2021-09-15 cs.CV cs.CL 79%

Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted as a long paper in the main conference of EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏