arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2411.17991 2025-11-25 cs.CV cs.CL 57%

VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format

VideoLLM 知何时言:通过视频-文本双人对话交互格式增强时间敏感视频理解

Yueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang, Jiansheng Wei, Huishuai Zhang, Dongyan Zhao

机构 * Wangxuan Institute of Computer Technology, Peking University(王炫计算机技术研究所,北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院) State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出视频-文本双人对话交互格式,通过MMDuetIT数据集和MAGQA任务提升VideoLLM在时间敏感任务中的表现,实现高效实时响应。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00091 2025-11-21 cs.CY cs.AI 57%

A First-Principles Based Risk Assessment Framework and the IEEE P3396 Standard

基于第一性原理的风险评估框架及IEEE P3396标准

Richard J. Tong, Marina Cortês, Jeanine A. DeFalco, Mark Underwood, Janusz Zalewski

机构 * Chair, IEEE Artificial Intelligence Standards Committee (AISC) Institute of Astrophysics Space Sciences, University of Lisbon, Portugal Vice Chair, IEEE Artificial Intelligence Standards Committee University of New Haven Florida Gulf Coast University, United States State Academy of Applied Sciences, Ciechanow, Poland

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出基于第一性原理的生成式AI风险评估框架,通过信息分类系统识别不同风险类型并归责于相关方,旨在提升AI治理的严谨性与责任性。

Comments 8 pages with 3 tables. This manuscript is prepared for publication by the Institute of Electrical and Electronics Engineers, Standards Association (IEEE-SA), Sponsor Committee - Artificial Intelligence Standards Committee (C/AISC) as a White Paper of Working Group p3396 at https://standards.ieee.org/ieee/3396/11379/

Journal ref 2025 IEEE Conference on Artificial Intelligence (CAI), pp. 1588-1595, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19082 2025-11-19 cs.CV 57%

Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference

Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11777 2025-11-19 cs.RO cs.CV 57%

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

Vinit Mehta, Charu Sharma, Karthick Thiyagarajan

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) SensR Lab Western Sydney University(Western Sydney University SensR实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 45 pages, 15 figures, MDPI Sensors Journal

Journal ref Sensors 2025, 25(20), 6394

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09030 2025-11-19 cs.LG 57%

Contextual Learning for Anomaly Detection in Tabular Data

Spencer King, Zhilu Zhang, Ruofan Yu, Baris Coskun, Wei Ding, Qian Cui

机构 * Amazon Web Services, Seattle, WA, USA(亚马逊网络服务,西雅图,WA,USA)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to TMLR. 26 pages, 4 figures, 8 tables, 1 algorithm, 8 datasets, contextual anomaly detection framework for tabular data

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13242 2025-11-18 cs.CV 57%

MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection

Junjie Wu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12404 2025-11-18 cs.MM cs.AI cs.SD 57%

SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs

Shail Desai, Aditya Pawar, Li Lin, Xin Wang, Shu Hu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12387 2025-11-18 cs.CL cs.AI 57%

From Phonemes to Meaning: Evaluating Large Language Models on Tamil

Jeyarajalingam Varsha, Menan Velayuthan, Sumirtha Karunakaran, Rasan Nivethiga, Kengatharaiyer Sarveswaran

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12020 2025-11-18 cs.CV 57%

LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension

Xianglong Shi, Silin Cheng, Sirui Zhao, Yunhan Jiang, Enhong Chen, Yang Liu, Sebastien Ourselin

机构 * University of Science and Technology of China(中国科学技术大学) The University of Hong Kong(香港大学) King’s College London(伦敦国王学院) Peking University(北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11212 2025-11-17 cs.CV 57%

MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI

Mohammad Areeb Qazi, Munachiso S Nwadike, Ibrahim Almakky, Mohammad Yaqub, Numan Saeed

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10923 2025-11-17 cs.CV 57%

Out-of-Distribution Detection with Positive and Negative Prompt Supervision Using Large Language Models

Zhixia He, Chen Zhao, Minglai Shao, Xintao Wu, Xujiang Zhao, Dong Li, Qin Tian, Linlin Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10914 2025-11-17 cs.CV 57%

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Mathematical Sciences, Fudan University(复旦大学数学学院) School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10518 2025-11-14 cs.CV cs.RO 57%

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

Wei Li, Renshan Zhang, Rui Shao, Zhijian Fang, Kaiwen Zhou, Zhuotao Tian, Liqiang Nie

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to AAAI 2026 (Oral), Project Page: https://github.com/JiuTian-VL/SemanticVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23740 2025-11-14 cs.CV cs.GR 57%

LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization

Ronghuan Wu, Wanchao Su, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Monash University(墨尔本大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Project Page: https://layerpeeler.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07983 2025-11-13 cs.CV 57%

ChexFract: From General to Specialized -- Enhancing Fracture Description Generation

Nikolay Nechaev, Evgeniia Przhezdzetskaia, Dmitry Umerenkov, Dmitry V. Dylov

机构 * Artificial Intelligence Research Institute (AIRI)(人工智能研究 institute (AIRI))

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10391 2025-11-13 cs.AI 57%

LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval

Yaoze Zhang, Rong Wu, Pinlong Cai, Xiaoman Wang, Guohang Yan, Song Mao, Ding Wang, Botian Shi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08018 2025-11-12 cs.CV 57%

High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection

Zhiyuan Chen, Yuelin Guo, Zitong Huang, Haoyu He, Renhao Lu, Weizhe Zhang

机构 * Institute of Cyberspace Security, Harbin Institute of Technology, Shenzhen(网络安全研究所,哈尔滨工业大学,深圳) Department of New Networks, Peng Cheng Laboratory(新网络部门,鹏城实验室) School of Cyberspace Science, Harbin Institute of Technology(空间科学学院,哈尔滨工业大学) Center on Machine Learning Research, Harbin Institute of Technology(机器学习研究中心,哈尔滨工业大学) Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments This work has been submitted to Pattern Recognition for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07982 2025-11-12 cs.CL cs.AI 57%

NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation

Maoqi Liu, Quan Fang, Yuhao Wu, Can Zhao, Yang Yang, Kaiquan Cai

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07966 2025-11-12 cs.CV 57%

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15267 2025-11-12 cs.CL cs.AI 57%

TraceCoder: Towards Traceable ICD Coding via Multi-Source Knowledge Integration

Mucheng Ren, He Chen, Yuchen Yan, Danqing Hu, Jun Xu, Xian Zeng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accpeted as BIBM 2025 Regular. 6 pages. Camera-Ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00598 2025-11-12 cs.CV 57%

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16227 2025-11-10 cs.CL cs.AI 57%

What Can String Probability Tell Us About Grammaticality?

Jennifer Hu, Ethan Gotlieb Wilcox, Siyuan Song, Kyle Mahowald, Roger P. Levy

机构 * Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University(哈佛大学自然与人工智能研究 institute)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17902 2025-11-10 cs.CV cs.CL 57%

TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection

Girish A. Koushik, Helen Treharne, Aditya Joshi, Diptesh Kanojia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to Special Track on AI for Social Impact (AISI) at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14245 2025-11-10 cs.CV cs.CL 57%

Towards Explainable Fake Image Detection with Multi-Modal Large Language Models

Yikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen, jun lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, Jianfu Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025; 14 pages including Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01083 2025-11-10 cs.RO cs.AI 57%

Affordance-based Robot Manipulation with Flow Matching

Fan Zhang, Michael Gienger

机构 * Honda Research Institute EU(本田欧洲研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04590 2025-11-07 cs.LG cs.IT math.IT 57%

Complexity as Advantage: A Regret-Based Perspective on Emergent Structure

Oshri Naparstek

机构 * IBM Research(IBM研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 15 pages. Under preparation for submission to ICML 2026. Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 57%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04847 2025-11-07 cs.CL cs.AI 57%

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) Vectara(Vectara公司) Iowa State University(爱荷华州立大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP Industry Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04112 2025-11-07 cs.CV 57%

SpatialLock: Precise Spatial Control in Text-to-Image Synthesis

Biao Liu, Yuanzhi Liang

机构 * The Sugon Group(神舟集团) TeleAI China Telecom(中国电信) Shanghai China(上海中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03549 2025-11-06 cs.SE cs.AI 57%

Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding

Ziv Nevo, Orna Raz, Karen Yorav

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 7 pages, 6 figures, to be published in AISM 2025, see https://aism25.github.io/aism25/

详情

展开后加载摘要…

URL PDF HTML 收藏