arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2503.08496 2025-03-12 cs.CV 57%

SuperCap: Multi-resolution Superpixel-based Image Captioning

Henry Senior, Luca Rossi, Gregory Slabaugh, Shanxin Yuan

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07456 2025-03-11 cs.CV 57%

Anatomy-Aware Conditional Image-Text Retrieval

Meng Zheng, Jiajin Zhang, Benjamin Planche, Zhongpai Gao, Terrence Chen, Ziyan Wu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06526 2025-03-11 cs.CV cs.MM 57%

TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos

Chen-Lin Zhang, Lin Sui, Shuming Liu, Fangzhou Mu, Zhangcheng Wang, Bernard Ghanem

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Code & models will be released at https://github.com/sming256/TimeLoc. The first 4 authors contributes equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06435 2025-03-11 cs.CV 57%

OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection

Adrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards, Krzysztof Czarnecki

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13654 2025-03-11 cs.CV 57%

GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting

Yuning Peng, Haiping Wang, Yuan Liu, Chenglu Wen, Zhen Dong, Bisheng Yang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Project page: https://pz0826.github.io/GAGS-Webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04533 2025-03-11 cs.CV 57%

Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation

Yongkang Li, Tianheng Cheng, Bin Feng, Wenyu Liu, Xinggang Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by CVPR 2025; Code & models: https://github.com/hustvl/MaskAdapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20548 2025-03-11 cs.RO cs.AI cs.HC 57%

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

Anxing Xiao, Nuwan Janaka, Tianrun Hu, Anshul Gupta, Kaixin Li, Cunjun Yu, David Hsu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

Comments Accepted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.20076 2025-03-11 cs.CV 57%

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Yuxuan Zhang, Tianheng Cheng, Lianghui Zhu, Rui Hu, Lei Liu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Preprint. Update: (1) better performance and (2) versatile segmentation. Code and models are available at: https://github.com/hustvl/EVF-SAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14491 2025-03-11 cs.AI 57%

Statistical Scenario Modelling and Lookalike Distributions for Multi-Variate AI Risk

Elija Perrier

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17859 2025-03-06 cs.CV cs.RO 57%

Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation

Yangxiao Lu, Jishnu Jaykumar P, Yunhui Guo, Nicholas Ruozzi, Yu Xiang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Project Page: https://irvlutd.github.io/NIDSNet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01151 2025-03-04 cs.CL cs.AI cs.IR 57%

ReaderLM-v2: Small Language Model for HTML to Markdown and JSON

Feng Wang, Zesheng Shi, Bo Wang, Nan Wang, Han Xiao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 9 pages, 10-12 refs

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00811 2025-03-04 cs.CV 57%

Evaluating and Predicting Distorted Human Body Parts for Generated Images

Lu Ma, Kaibo Cao, Hao Liang, Jiaxin Lin, Zhuang Li, Yuhong Liu, Jihong Zhang, Wentao Zhang, Bin Cui

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01345 2025-03-04 cs.RO cs.CV 57%

Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy

Ricardo Garcia, Shizhe Chen, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20724 2025-03-04 cs.CL cs.IR cs.LG 57%

Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation

Mufei Li, Siqi Miao, Pan Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted by ICLR 2025; Code available at https://github.com/Graph-COM/SubgraphRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12998 2025-03-04 eess.AS cs.AI cs.CL cs.SD 57%

Coding Speech through Vocal Tract Kinematics

Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal, Gopala K. Anumanchipalli

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 18, no. 8, pp. 1427-1440, Dec. 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20669 2025-03-03 cs.CV 57%

EndoPBR: Material and Lighting Estimation for Photorealistic Surgical Simulations via Physically-based Rendering

John J. Han, Jie Ying Wu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19782 2025-02-28 cs.CV 57%

Open-Vocabulary Semantic Part Segmentation of 3D Human

Keito Suzuki, Bang Du, Girish Krishnan, Kunyao Chen, Runfa Blark Li, Truong Nguyen

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 3DV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15278 2025-02-28 cs.CV 57%

PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions

Weifeng Lin, Xinyu Wei, Renrui Zhang, Le Zhuo, Shitian Zhao, Siyuan Huang, Huan Teng, Junlin Xie, Yu Qiao, Peng Gao, Hongsheng Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Code is released at https://github.com/AFeng-x/PixWizard

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08772 2025-02-28 cs.CV cs.CL 57%

MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs

Xuannan Liu, Zekun Li, Peipei Li, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICLR 2025, Project page: https://liuxuannan.github.io/MMFakeBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10003 2025-02-27 cs.AI cs.DL 57%

Towards a Knowledge Graph for Models and Algorithms in Applied Mathematics

Björn Schembera, Frank Wübbeling, Hendrik Kleikamp, Burkhard Schmidt, Aurela Shehu, Marco Reidelbach, Christine Biedinger, Jochen Fiedler, Thomas Koprucki, Dorothea Iglezakis, Dominik Göddeke

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Preprint submitted to the 18th International Conference on Metadata and Semantics Research 2024 and published as a full, revised article

Journal ref Sfakakis, M., Garoufallou, E., Damigos, M., Salaba, A., Papatheodorou, C. (eds) Metadata and Semantic Research. MTSR 2024. Communications in Computer and Information Science, vol 2331. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18736 2025-02-27 cs.HC cs.AI 57%

AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools

Nathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown, Hugo Romat, Michel Pahud, Nicolai Marquardt, Kori Inkpen, Ken Hinckley

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 18 pages, 10 figures. To appear in the Proceedings of the 2025 ACM CHI Conference on Human Factors in Computing Systems, Yokohama, Japan. https://hugoromat.github.io/ai_instruments/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17763 2025-02-26 cs.CR cs.AI cs.DC cs.PF 57%

Design and implementation of a distributed security threat detection system integrating federated learning and multimodal LLM

Yuqing Wang, Xiao Yang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15872 2025-02-25 cs.CL cs.AI cs.SE 57%

MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use

Zaid Khan, Ali Farhadi, Ranjay Krishna, Luca Weihs, Mohit Bansal, Tanmay Gupta

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Project page: zaidkhan.me/MutaGReP

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15124 2025-02-24 math.NA cs.LG cs.NA math.DG 57%

Curvature Corrected Nonnegative Manifold Data Factorization

Joyce Chew, Willem Diepeveen, Deanna Needell

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14994 2025-02-24 cs.CV 57%

LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection

Qingyuan Liu, Yun-Yun Tsai, Ruijian Zha, Victoria Li, Pengyuan Shi, Chengzhi Mao, Junfeng Yang

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13743 2025-02-20 cs.AI 57%

Inference of Abstraction for Grounded Predicate Logic

Hiroyuki Kido

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13161 2025-02-20 q-bio.NC cs.AI 57%

Noumenal Labs White Paper: How To Build A Brain

Maxwell J. D. Ramstead, Candice Pattisapu, Jason Fox, Jeff Beck

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18295 2025-02-19 cs.CV 57%

Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention

Weitai Kang, Mengxue Qu, Jyoti Kini, Yunchao Wei, Mubarak Shah, Yan Yan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11891 2025-02-18 cs.CV 57%

From Open-Vocabulary to Vocabulary-Free Semantic Segmentation

Klara Reichard, Giulia Rizzoli, Stefano Gasperini, Lukas Hoyer, Pietro Zanuttigh, Nassir Navab, Federico Tombari

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Submitted to: Pattern Recognition Letters, Klara Reichard and Giulia Rizzoli equally contributed to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10675 2025-02-18 cs.CV 57%

Hierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language Model

Weilin Sun, Xinran Li, Manyi Li, Kai Xu, Xiangxu Meng, Lei Meng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏