arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2310.08840 2023-10-16 cs.CL cs.AI 57%

Large Language Models as Source Planner for Personalized Knowledge-grounded Dialogue

Hongru Wang, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, Kam-Fai Wong

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01470 2023-10-04 cs.AI 57%

Challenges in Modelling and Solving Plotting with PDDL

Joan Espasa, Ian Miguel, Peter Nightingale, András Z. Salamon, Mateu Villaret

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2110.14397

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00240 2023-10-03 cs.CV eess.IV 57%

Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

Siyu Jiao, Yunchao Wei, Yaowei Wang, Yao Zhao, Humphrey Shi

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16649 2023-09-29 cs.CV 57%

FLIP: Cross-domain Face Anti-spoofing with Language Guidance

Koushik Srivatsan, Muzammal Naseer, Karthik Nandakumar

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to ICCV-2023. Project Page: https://koushiksrivats.github.io/FLIP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14084 2023-09-26 cs.CL cs.AI cs.IR 57%

Comprehensive Overview of Named Entity Recognition: Models, Domain-Specific Applications and Challenges

Kalyani Pakhale

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13247 2023-09-26 cs.CV 57%

Multi-modal Domain Adaptation for REG via Relation Transfer

Yifan Ding, Liqiang Wang, Boqing Gong

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09456 2023-09-19 cs.CV 57%

Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection

Chenming Zhu, Wenwei Zhang, Tai Wang, Xihui Liu, Kai Chen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 17 pages, 7 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08827 2023-09-19 cs.CL cs.AI 57%

S3-DST: Structured Open-Domain Dialogue Segmentation and State Tracking in the Era of LLMs

Sarkar Snigdha Sarathi Das, Chirag Shah, Mengting Wan, Jennifer Neville, Longqi Yang, Reid Andersen, Georg Buscher, Tara Safavi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06581 2023-09-14 cs.CV 57%

Zero-Shot Visual Classification with Guided Cropping

Piyapat Saranrittichai, Mauricio Munoz, Volker Fischer, Chaithanya Kumar Mummadi

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04634 2023-09-13 cs.CL cs.AI 57%

Open-world Story Generation with Structured Knowledge Enhancement: A Comprehensive Survey

Yuxin Wang, Jieru Lin, Zhiwei Yu, Wei Hu, Börje F. Karlsson

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted in Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07935 2023-09-12 cs.AI cs.LO cs.RO 57%

Strong-AI Autoepistemic Robots Build on Intensional First Order Logic

Zoran Majkic

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 25 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03874 2023-09-08 cs.CV 57%

Box-based Refinement for Weakly Supervised and Unsupervised Localization Tasks

Eyal Gomel, Tal Shaharabany, Lior Wolf

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03483 2023-09-08 cs.CV 57%

DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners

Clarence Lee, M Ganesh Kumar, Cheston Tan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01151 2023-09-06 cs.CV 57%

EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment

Cheng Shi, Sibei Yang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2023; Project Page: https://chengshiest.github.io/edadet

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00227 2023-09-04 cs.CV 57%

What Makes Good Open-Vocabulary Detector: A Disassembling Perspective

Jincheng Li, Chunyu Xie, Xiaoyu Wu, Bin Wang, Dawei Leng

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Journal ref KDD workshop 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16649 2023-09-01 cs.CV 57%

Learning with Multi-modal Gradient Attention for Explainable Composed Image Retrieval

Prateksha Udhayanan, Srikrishna Karanam, Balaji Vasan Srinivasan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15962 2023-09-01 cs.RO cs.AI 57%

WALL-E: Embodied Robotic WAiter Load Lifting with Large Language Model

Tianyu Wang, Yifan Li, Haitao Lin, Xiangyang Xue, Yanwei Fu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 14 pages, 8 figures. See https://star-uu-wang.github.io/WALL-E/

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06712 2023-08-21 cs.CV 57%

What does CLIP know about a red circle? Visual prompt engineering for VLMs

Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2023 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07918 2023-08-16 cs.CV 57%

Helping Hands: An Object-Aware Ego-Centric Video Recognition Model

Chuhan Zhang, Ankush Gupta, Andrew Zisserman

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05221 2023-08-11 cs.CV 57%

Open-vocabulary Object Segmentation with Diffusion Models

Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00849 2023-08-01 cs.CV 57%

Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection

Yanxin Long, Jianhua Han, Runhui Huang, Xu Hang, Yi Zhu, Chunjing Xu, Xiaodan Liang

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12935 2023-07-25 cs.CL cs.AI 57%

Rule By Example: Harnessing Logical Rules for Explainable Hate Speech Detection

Christopher Clarke, Matthew Hall, Gaurav Mittal, Ye Yu, Sandra Sajeev, Jason Mars, Mei Chen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments ACL 2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11435 2023-07-25 eess.AS cs.AI cs.CL cs.SD 57%

Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model

Puyuan Peng, Shang-Wen Li, Okko Räsänen, Abdelrahman Mohamed, David Harwath

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Interspeech 2023. Code & Model: https://github.com/jasonppy/syllable-discovery

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05303 2023-07-25 cs.CV cs.CL 57%

ELVIS: Empowering Locality of Vision Language Pre-training with Intra-modal Similarity

Sumin Seo, JaeWoong Shin, Jaewoo Kang, Tae Soo Kim, Thijs Kooi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09484 2023-07-24 q-bio.BM cs.CE cs.LG physics.chem-ph 57%

MolFM: A Multimodal Molecular Foundation Model

Yizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu, Zaiqing Nie

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 31 pages, 15 figures, and 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11723 2023-07-24 eess.AS cs.AI cs.SD 57%

Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech

Cheol Jun Cho, Peter Wu, Abdelrahman Mohamed, Gopala K. Anumanchipalli

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10226 2023-07-21 cs.AI 57%

On Loop Formulas with Variables

Joohyung Lee, Yunsong Meng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 10 pages. In Proc. Eleventh International Conference on Principles of Knowledge Representation and Reasoning (KR 2008), pages 444-453. arXiv admin note: text overlap with arXiv:1401.3898

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10225 2023-07-21 cs.AI cs.SC 57%

First-Order Stable Model Semantics with Intensional Functions

Michael Bartholomew, Joohyung Lee

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 69 pages

Journal ref Artificial Intelligence 273, 56-93, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09756 2023-07-20 cs.CV 57%

Generative Prompt Model for Weakly Supervised Object Localization

Yuzhong Zhao, Qixiang Ye, Weijia Wu, Chunhua Shen, Fang Wan

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Journal ref International Conference on Computer Vision Conference (ICCV2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09184 2023-07-19 cs.CV 57%

You've Got Two Teachers: Co-evolutionary Image and Report Distillation for Semi-supervised Anatomical Abnormality Detection in Chest X-ray

Jinghan Sun, Dong Wei, Zhe Xu, Donghuan Lu, Hong Liu, Liansheng Wang, Yefeng Zheng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏