arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2209.09066 2023-01-10 cs.AI 57%

Specifying and Exploiting Non-Monotonic Domain-Specific Declarative Heuristics in Answer Set Programming

Richard Comploi-Taupe, Gerhard Friedrich, Konstantin Schekotihin, Antonius Weinzierl

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05636 2023-01-06 cs.CV 57%

VL-NMS: Breaking Proposal Bottlenecks in Two-Stage Visual-Language Matching

Chenchi Zhang, Wenbo Ma, Jun Xiao, Hanwang Zhang, Jian Shao, Yueting Zhuang, Long Chen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2009.01449

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16208 2022-12-09 cs.CV 57%

SLAN: Self-Locator Aided Network for Cross-Modal Understanding

Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen, Jiang-Jiang Liu, Bo Ren, Ming-Ming Cheng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09817 2022-12-08 cs.CV cs.CL 57%

Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing

Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C. Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, Hoifung Poon, Ozan Oktay

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments To appear in ECCV 2022. Code: https://aka.ms/biovil-code Dataset: https://aka.ms/ms-cxr Demo Notebook: https://aka.ms/biovil-demo-notebook

Journal ref Computer Vision - ECCV 2022, LNCS vol 13696, pp 1-21

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14843 2022-11-29 cs.CV 57%

Learning Object-Language Alignments for Open-Vocabulary Object Detection

Chuang Lin, Peize Sun, Yi Jiang, Ping Luo, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan, Jianfei Cai

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14347 2022-11-29 cs.LG 57%

The smooth output assumption, and why deep networks are better than wide ones

Luis Sa-Couto, Jose Miguel Ramos, Andreas Wichert

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11134 2022-11-28 cs.CV 57%

Open Vocabulary Object Detection with Proposal Mining and Prediction Equalization

Peixian Chen, Kekai Sheng, Mengdan Zhang, Mingbao Lin, Yunhang Shen, Shaohui Lin, Bo Ren, Ke Li

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15867 2022-11-21 cs.CV cs.CL 57%

Image Retrieval from Contextual Descriptions

Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, Siva Reddy

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments accepted to ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08704 2022-11-17 cs.CV 57%

A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge

Sicheng Mo, Fangzhou Mu, Yin Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04534 2022-11-10 cs.CV cs.CL 57%

Going for GOAL: A Resource for Grounded Football Commentaries

Alessandro Suglia, José Lopes, Emanuele Bastianelli, Andrea Vanzo, Shubham Agarwal, Malvina Nikandrou, Lu Yu, Ioannis Konstas, Verena Rieser

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Preprint formatted using the ACM Multimedia template (8 pages + appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09862 2022-10-26 cs.RO cs.AI 57%

Learning to Act with Affordance-Aware Multimodal Neural SLAM

Zhiwei Jia, Kaixiang Lin, Yizhou Zhao, Qiaozi Gao, Govind Thattai, Gaurav Sukhatme

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12997 2022-10-25 cs.CL cs.CV 57%

Are Current Decoding Strategies Capable of Facing the Challenges of Visual Dialogue?

Amit Kumar Chaudhary, Alex J. Lucassen, Ioanna Tsani, Alberto Testoni

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted at INLG 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12687 2022-10-25 cs.CL cs.AI 57%

BotsTalk: Machine-sourced Framework for Automatic Curation of Large-scale Multi-skill Dialogue Datasets

Minju Kim, Chaehyeong Kim, Yongho Song, Seung-won Hwang, Jinyoung Yeo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted to EMNLP2022. Code and data are available at https://github.com/convei-lab/BotsTalk

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12565 2022-10-25 cs.CL cs.LG 57%

A Visual Tour Of Current Challenges In Multimodal Language Models

Shashank Sonkar, Naiming Liu, Richard G. Baraniuk

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10828 2022-10-21 cs.CV 57%

Grounded Video Situation Recognition

Zeeshan Khan, C. V. Jawahar, Makarand Tapaswi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2022. Project Page: https://zeeshank95.github.io/grvidsitu

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09407 2022-10-18 cs.CV 57%

DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, Hang Xu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03825 2022-10-11 cs.AI cs.RO 57%

See, Plan, Predict: Language-guided Cognitive Planning with Video Prediction

Maria Attarian, Advaya Gupta, Ziyi Zhou, Wei Yu, Igor Gilitschenski, Animesh Garg

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03416 2022-10-10 cs.CV 57%

Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding

Constantin Seibold, Simon Reiß, Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon Sung Kim, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 33rd British Machine Vision Conference (BMVC 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02667 2022-10-07 cs.AI cs.CY 57%

A Human Rights-Based Approach to Responsible AI

Vinodkumar Prabhakaran, Margaret Mitchell, Timnit Gebru, Iason Gabriel

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Presented as a (non-archival) poster at the 2022 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization or (EAAMO '22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02365 2022-10-06 cs.CV 57%

SoccerNet 2022 Challenges Results

Silvio Giancola, Anthony Cioppa, Adrien Deliège, Floriane Magera, Vladimir Somers, Le Kang, Xin Zhou, Olivier Barnich, Christophe De Vleeschouwer, Alexandre Alahi, Bernard Ghanem, Marc Van Droogenbroeck, Abdulrahman Darwish, Adrien Maglo, Albert Clapés, Andreas Luyts, Andrei Boiarov, Artur Xarles, Astrid Orcesi, Avijit Shah, Baoyu Fan, Bharath Comandur, Chen Chen, Chen Zhang, Chen Zhao, Chengzhi Lin, Cheuk-Yiu Chan, Chun Chuen Hui, Dengjie Li, Fan Yang, Fan Liang, Fang Da, Feng Yan, Fufu Yu, Guanshuo Wang, H. Anthony Chan, He Zhu, Hongwei Kan, Jiaming Chu, Jianming Hu, Jianyang Gu, Jin Chen, João V. B. Soares, Jonas Theiner, Jorge De Corte, José Henrique Brito, Jun Zhang, Junjie Li, Junwei Liang, Leqi Shen, Lin Ma, Lingchi Chen, Miguel Santos Marques, Mike Azatov, Nikita Kasatkin, Ning Wang, Qiong Jia, Quoc Cuong Pham, Ralph Ewerth, Ran Song, Rengang Li, Rikke Gade, Ruben Debien, Runze Zhang, Sangrok Lee, Sergio Escalera, Shan Jiang, Shigeyuki Odashima, Shimin Chen, Shoichi Masui, Shouhong Ding, Sin-wai Chan, Siyu Chen, Tallal El-Shabrawy, Tao He, Thomas B. Moeslund, Wan-Chi Siu, Wei Zhang, Wei Li, Xiangwei Wang, Xiao Tan, Xiaochuan Li, Xiaolin Wei, Xiaoqing Ye, Xing Liu, Xinying Wang, Yandong Guo, Yaqian Zhao, Yi Yu, Yingying Li, Yue He, Yujie Zhong, Zhenhua Guo, Zhiheng Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted at ACM MMSports 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02515 2022-09-20 cs.CV 57%

Fine-Grained Semantically Aligned Vision-Language Pre-Training

Juncheng Li, Xin He, Longhui Wei, Long Qian, Linchao Zhu, Lingxi Xie, Yueting Zhuang, Qi Tian, Siliang Tang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10400 2022-08-18 cs.CV 57%

Correspondence Matters for Video Referring Expression Comprehension

Meng Cao, Ji Jiang, Long Chen, Yuexian Zou

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07852 2022-08-17 cs.CL cs.HC cs.LG 57%

Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models

Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, Alexander M. Rush

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 9 pages content, 2 pages references

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.17273 2022-08-10 cs.CV 57%

FindIt: Generalized Localization with Natural Language Queries

Weicheng Kuo, Fred Bertsch, Wei Li, AJ Piergiovanni, Mohammad Saffar, Anelia Angelova

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to ECCV 2022 (European Conference on Computer Vision)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01534 2022-08-09 cs.IR cs.AI cs.HC 57%

Towards Psychologically-Grounded Dynamic Preference Models

Mihaela Curmei, Andreas Haupt, Dylan Hadfield-Menell, Benjamin Recht

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments In Sixteenth ACM Conference on Recommender Systems, September 18-23, 2022, Seattle, WA, USA, 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14205 2022-07-29 cs.RO cs.AI 57%

DoRO: Disambiguation of referred object for embodied agents

Pradip Pramanick, Chayan Sarkar, Sayan Paul, Ruddra dev Roychoudhury, Brojeshwar Bhowmick

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted in IEEE Robotics & Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07011 2022-07-25 cs.CL cs.AI 57%

Towards Socially Intelligent Agents with Mental State Transition and Human Utility

Liang Qiu, Yizhou Zhao, Yuan Liang, Pan Lu, Weiyan Shi, Zhou Yu, Song-Chun Zhu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Long paper accepted by SIGDIAL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10362 2022-07-22 cs.CV 57%

LocVTP: Video-Text Pre-training for Temporal Localization

Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, Yuexian Zou

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.08212 2022-07-19 cs.CL cs.AI 57%

RT-KGD: Relation Transition Aware Knowledge-Grounded Dialogue Generation

Kexin Wang, Zhixu Li, Jiaan Wang, Jianfeng Qu, Ying He, An Liu, Lei Zhao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments ISWC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08964 2022-07-15 cs.CL cs.AI 57%

A recipe for annotating grounded clarifications

Luciana Benotti, Patrick Blackburn

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted for publication at the 2021 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2021)

Journal ref https://aclanthology.org/2021.naacl-main.320

详情

展开后加载摘要…

URL PDF HTML 收藏