arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2501.07515 2025-01-14 cs.NE cs.AI 57%

The Paradox of Success in Evolutionary and Bioinspired Optimization: Revisiting Critical Issues, Key Studies, and Methodological Pathways

Daniel Molina, Javier Del Ser, Javier Poyatos, Francisco Herrera

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 38 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06761 2025-01-14 cs.CV 57%

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning

Ji Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park, Hyunwoo J. Kim

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17114 2025-01-14 cs.AI cs.ET 57%

Decentralized Governance of Autonomous AI Agents

Tomer Jordi Chaffer, Charles von Goins, Bayo Okusanya, Dontrail Cotlage, Justin Goldston

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04247 2025-01-09 cs.CV 57%

3D Part Segmentation via Geometric Aggregation of 2D Visual Features

Marco Garosi, Riccardo Tedoldi, Davide Boscaini, Massimiliano Mancini, Nicu Sebe, Fabio Poiesi

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Published in WACV 2025. Project page: https://3d-cops.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16496 2025-01-08 cs.LG 57%

Can Out-of-Domain data help to Learn Domain-Specific Prompts for Multimodal Misinformation Detection?

Amartya Bhattacharya, Debarshi Brahma, Suraj Nagaje Mahadev, Anmol Asati, Vikas Verma, Soma Biswas

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17325 2025-01-07 cs.CV 57%

Feature Based Methods in Domain Adaptation for Object Detection: A Review Paper

Helia Mohamadi, Mohammad Ali Keyvanrad, Mohammad Reza Mohammadi

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments 48 pages, 13 figures, It will be submitted to the Artificial Intelligence Review journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05130 2025-01-07 cs.AI 57%

From Chain to Tree: Refining Chain-like Rules into Tree-like Rules on Knowledge Graphs

Wangtao Sun, Shizhu He, Jun Zhao, Kang Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19047 2025-01-07 cs.LG math.DS 57%

Theoretical Foundations of Deep Selective State-Space Models

Nicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi, Terry Lyons

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Fina NeurIPS Camera Ready Version w/ minor edits

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00877 2025-01-06 cs.CV 57%

FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation

Bingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05184 2025-01-03 cs.CV 57%

The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better

Scott Geng, Cheng-Yu Hsieh, Vivek Ramanujan, Matthew Wallingford, Chun-Liang Li, Pang Wei Koh, Ranjay Krishna

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Correspondence to sgeng at cs dot washington dot edu. RK and PWK equally advised the project

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19944 2024-12-31 cs.CV 57%

Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark

Lukas Picek, Vojtěch Čermák, Marek Hanzl

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19648 2024-12-30 cs.CV cs.MM 57%

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

X. Feng, D. Zhang, S. Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ICASSP '25 ! Code: https://github.com/XiaokunFeng/CTVLT

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19492 2024-12-30 cs.CV cs.MM 57%

Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation

Chengyang Ye, Yunzhi Zhuge, Pingping Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19155 2024-12-30 cs.CV cs.CL 57%

Referencing Where to Focus: Improving VisualGrounding with Referential Query

Yabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin, Sanping Zhou, Ming Yang, Le Wang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by NIPS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18672 2024-12-30 cs.CL cs.AI 57%

From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs

Ratnesh Kumar Joshi, Sagnik Sengupta, Asif Ekbal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 14 Pages, 5 Tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17800 2024-12-24 cs.CV 57%

Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection

Yitong Chen, Wenhao Yao, Lingchen Meng, Sihong Wu, Zuxuan Wu, Yu-Gang Jiang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Code is available at https://github.com/Row11n/Prova/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14006 2024-12-19 cs.CV 57%

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, Yujiu Yang

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06286 2024-12-19 cs.CV 57%

No Annotations for Object Detection in Art through Stable Diffusion

Patrick Ramos, Nicolas Gonthier, Selina Khan, Yuta Nakashima, Noa Garcia

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 8 pages, 6 figures, to be published in WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06268 2024-12-19 cs.CV 57%

Open-Vocabulary High-Resolution 3D (OVHR3D) Data Segmentation and Annotation Framework

Jiuyi Xu, Meida Chen, Andrew Feng, Zifan Yu, Yangming Shi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Journal ref Interservice/Industry Training, Simulation and Education Conference (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11481 2024-12-19 cs.CV cs.RO 57%

Exploring Emerging Trends and Research Opportunities in Visual Place Recognition

Antonios Gasteratos, Konstantinos A. Tsintotas, Tobias Fischer, Yiannis Aloimonos, Michael Milford

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 2 pages, 1 figure. 40th Anniversary of the IEEE Conference on Robotics and Automation (ICRA@40), Rotterdam, Netherlands, September 23-26, 2024

Journal ref 40th Anniversary of the IEEE Conference on Robotics and Automation (ICRA@40), Rotterdam, Netherlands, September 23-26, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12798 2024-12-18 cs.CV 57%

ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation

Shiqi Huang, Shuting He, Bihan Wen

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments AAAI 2025, code see https://github.com/HuangShiqi128/ZoRI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12766 2024-12-18 cs.CV 57%

Towards a Training Free Approach for 3D Scene Editing

Vivek Madhavaram, Shivangana Rawat, Chaitanya Devaguptapu, Charu Sharma, Manohar Kaul

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12683 2024-12-18 cs.CV 57%

ShiftedBronzes: Benchmarking and Analysis of Domain Fine-Grained Classification in Open-World Settings

Rixin Zhou, Honglin Pang, Qian Zhang, Ruihua Qi, Xi Yang, Chuntao Li

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments 9pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12675 2024-12-18 cs.CV 57%

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries

Wangyu Xue, Chen Qian, Jiayi Wu, Yang Zhou, Wentao Liu, Ju Ren, Siming Fan, Yaoxue Zhang

专题命中 视觉定位与Grounding :InternVL(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01002 2024-12-17 cs.CL cs.AI 57%

Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries

Zelalem Gero, Chandan Singh, Yiqing Xie, Sheng Zhang, Praveen Subramanian, Paul Vozila, Tristan Naumann, Jianfeng Gao, Hoifung Poon

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Published in ML4H Findings 2024, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10380 2024-12-17 cs.HC cs.AI 57%

Challenges in Human-Agent Communication

Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, Daniel S. Weld

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07942 2024-12-12 cs.LG cond-mat.dis-nn 57%

Neural Scaling Laws Rooted in the Data Distribution

Ari Brill

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 24 pages, 10 figures, Code: https://github.com/aribrill/scaling-laws-paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06071 2024-12-10 cs.CV 57%

GlocalCLIP: Object-agnostic Global-Local Prompt Learning for Zero-shot Anomaly Detection

Jiyul Ham, Yonggon Jung, Jun-Geol Baek

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 29 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04441 2024-12-06 cs.CV 57%

Learning Artistic Signatures: Symmetry Discovery and Style Transfer

Emma Finn, T. Anderson Keller, Emmanouil Theodosis, Demba E. Ba

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01551 2024-12-06 cs.CV 57%

ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection

Maryam Hosseini, Marco Cipriano, Sedigheh Eslami, Daniel Hodczak, Liu Liu, Andres Sevtsuk, Gerard de Melo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏