arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2309.02427 2024-03-18 cs.AI cs.CL cs.LG cs.SC 62%

Cognitive Architectures for Language Agents

Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, Thomas L. Griffiths

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments v3 is TMLR camera ready version. 19 pages of main content, 5 figures. The first two authors contributed equally, order decided by coin flip. A CoALA-based repo of recent work on language agents: https://github.com/ysymyth/awesome-language-agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03179 2024-03-15 cs.CV cs.LG 62%

SLiMe: Segment Like Me

Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi Amiri, Ghassan Hamarneh

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08426 2024-03-14 cs.CV cs.AI 62%

Language-Driven Visual Consensus for Zero-Shot Semantic Segmentation

Zicheng Zhang, Tong Zhang, Yi Zhu, Jianzhuang Liu, Xiaodan Liang, QiXiang Ye, Wei Ke

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00272 2024-03-06 cs.LG cs.CV q-bio.TO 62%

A Survey on Graph-Based Deep Learning for Computational Histopathology

David Ahmedt-Aristizabal, Mohammad Ali Armin, Simon Denman, Clinton Fookes, Lars Petersson

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Preprint submitted to Computerized Medical Imaging and Graphics

Journal ref Volume 95, January 2022, 102027

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18127 2024-03-01 cs.LG cs.AI cs.CL 62%

Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models

Xue Yan, Yan Song, Xinyu Cui, Filippos Christianos, Haifeng Zhang, David Henry Mguni, Jun Wang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12846 2024-02-21 cs.CV cs.AI 62%

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments AAAI 2024. Project page at https://limirs.github.io/ConVQG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11984 2024-02-20 cs.NE cs.AI cs.LG 62%

Hebbian Learning based Orthogonal Projection for Continual Learning of Spiking Neural Networks

Mingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He, Zhouchen Lin

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted by ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13839 2024-02-19 cs.CV cs.LG 62%

Q-SENN: Quantized Self-Explaining Neural Networks

Thomas Norrenbrock, Marco Rudolph, Bodo Rosenhahn

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted to AAAI 2024, SRRAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00319 2024-02-02 cs.CV cs.AI 62%

SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling

Eileen Wang, Soyeon Caren Han, Josiah Poon

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01361 2024-01-23 cs.LG cs.CL cs.CV cs.RO 62%

GenSim: Generating Robotic Simulation Tasks via Large Language Models

Lirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar, Chen Bao, Yuzhe Qin, Bailin Wang, Huazhe Xu, Xiaolong Wang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments See our project website (https://liruiw.github.io/gensim), demo and datasets (https://huggingface.co/spaces/Gen-Sim/Gen-Sim), and code (https://github.com/liruiw/GenSim) for more details

Journal ref International Conference on Learning Representations (ICLR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08973 2024-01-18 cs.CV cs.AI cs.CL 62%

OCTO+: A Suite for Automatic Open-Vocabulary Object Placement in Mixed Reality

Aditya Sharma, Luke Yoffe, Tobias Höllerer

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 2024 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIXVR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07883 2024-01-17 cs.LG cs.AI cs.CL cs.IR 62%

The Chronicles of RAG: The Retriever, the Chunk and the Generator

Paulo Finardi, Leonardo Avila, Rodrigo Castaldoni, Pedro Gengo, Celio Larcher, Marcos Piau, Pablo Costa, Vinicius Caridá

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 15 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05335 2024-01-11 cs.CV cs.GR cs.LG 62%

InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes

Mohamad Shahbazi, Liesbeth Claessens, Michael Niemeyer, Edo Collins, Alessio Tonioni, Luc Van Gool, Federico Tombari

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03910 2024-01-09 cs.CL cs.AI cs.LG 62%

A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates

Raphaël Millière, Cameron Buckner

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01364 2024-01-04 q-bio.NC cs.AI cs.LG cs.NE 62%

Multi-Modal Cognitive Maps based on Neural Networks trained on Successor Representations

Paul Stoewer, Achim Schilling, Andreas Maier, Patrick Krauss

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18341 2023-12-27 cs.PL cs.AI cs.LG 62%

Coarse-Tuning Models of Code with Reinforcement Learning Feedback

Abhinav Jain, Chima Adiole, Swarat Chaudhuri, Thomas Reps, Chris Jermaine

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12815 2023-12-21 cs.CV cs.AI cs.CL 62%

OCTOPUS: Open-vocabulary Content Tracking and Object Placement Using Semantic Understanding in Mixed Reality

Luke Yoffe, Aditya Sharma, Tobias Höllerer

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments IEEE International Symposium on Mixed and Augmented Reality (ISMAR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04498 2023-12-19 cs.CV cs.AI cs.CL 62%

NExT-Chat: An LMM for Chat, Detection and Segmentation

Ao Zhang, Yuan Yao, Wei Ji, Zhiyuan Liu, Tat-Seng Chua

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Technical Report (https://next-chatv.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13550 2023-12-12 cs.CV cs.AI 62%

I-AI: A Controllable & Interpretable AI System for Decoding Radiologists' Intense Focus for Accurate CXR Diagnoses

Trong Thang Pham, Jacob Brecheisen, Anh Nguyen, Hien Nguyen, Ngan Le

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17390 2023-12-07 cs.CL cs.AI cs.LG cs.MA cs.RO 62%

SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks

Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, Xiang Ren

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2023 (spotlight). Project website: https://swiftsage.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16464 2023-11-29 cs.CV cs.AI 62%

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

Yicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma, Hengwei Bian, Yatai Ji, Yujiu Yang, Xiu Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07871 2023-11-27 cs.AI cs.LG 62%

The SocialAI School: Insights from Developmental Psychology Towards Artificial Socio-Cultural Agents

Grgur Kovač, Rémy Portelas, Peter Ford Dominey, Pierre-Yves Oudeyer

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Preprint, see v1 for a shorter version (accepted at the "Workshop on Theory-of-Mind" at ICML 2023) See project website for demo and code: https://sites.google.com/view/socialai-school

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12827 2023-11-22 cs.LG cs.CV 62%

Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models

Guillermo Ortiz-Jimenez, Alessandro Favero, Pascal Frossard

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Journal ref Advances in Neural Information Processing Systems 36 (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11754 2023-11-21 cs.CV cs.AI 62%

A Large-Scale Car Parts (LSCP) Dataset for Lightweight Fine-Grained Detection

Wang Jie, Zhong Yilin, Cao Qianqian

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19909 2023-11-21 cs.CV cs.LG 62%

Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks

Micah Goldblum, Hossein Souri, Renkun Ni, Manli Shu, Viraj Prabhu, Gowthami Somepalli, Prithvijit Chattopadhyay, Mark Ibrahim, Adrien Bardes, Judy Hoffman, Rama Chellappa, Andrew Gordon Wilson, Tom Goldstein

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03053 2023-11-07 cs.CV cs.AI 62%

Masking Hyperspectral Imaging Data with Pretrained Models

Elias Arbash, Andréa de Lima Ribeiro, Sam Thiele, Nina Gnann, Behnood Rasti, Margret Fuchs, Pedram Ghamisi, Richard Gloaguen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00206 2023-11-02 cs.CV cs.AI 62%

ChatGPT-Powered Hierarchical Comparisons for Image Classification

Zhiyuan Ren, Yiyang Su, Xiaoming Liu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Neurips 2023 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04704 2023-10-10 cs.CV cs.AI cs.CL 62%

Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition

Shuhuai Ren, Aston Zhang, Yi Zhu, Shuai Zhang, Shuai Zheng, Mu Li, Alex Smola, Xu Sun

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Code is available at https://github.com/amazon-science/prompt-pretraining

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03779 2023-10-09 cs.AI cs.CL cs.LG cs.RO 62%

HandMeThat: Human-Robot Communication in Physical and Social Environments

Yanming Wan, Jiayuan Mao, Joshua B. Tenenbaum

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2022 (Dataset and Benchmark Track). First two authors contributed equally. Project page: http://handmethat.csail.mit.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00166 2023-10-03 cs.AI cs.LG 62%

Motif: Intrinsic Motivation from Artificial Intelligence Feedback

Martin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, Mikael Henaff

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments The first two authors equally contributed - order decided by coin flip

详情

展开后加载摘要…

URL PDF HTML 收藏