arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-07 至 2025-10-07 共收录 22 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 22 篇

2510.04477 2025-10-07 cs.CV cs.AI cs.CL cs.LG 87%

MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models

Soo Yong Kim, Suin Cho, Vincent-Daniel Yun, Gyeongyeon Hwang

机构 * A.I.MATICS Inc(A.I.MATICS公司) Boston University(波士顿大学) University of Southern California(南加州大学) Heuron(Heuron公司) MODULABS, Open Neural Networks Research Lab(MODULABS,开放神经网络研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12266 2025-10-07 cs.CV cs.AI cs.CL 84%

CBVLM: Training-free Explainable Concept-based Large Vision Language Models for Medical Image Classification

Cristiano Patrício, Isabel Rio-Torto, Jaime S. Cardoso, Luís F. Teixeira, João C. Neves

机构 * INESC TEC NOVA LINCS Universidade da Beira Interior(贝拉蒙特大学) Universidade do Porto(波尔图大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in Computers in Biology and Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04039 2025-10-07 cs.CV cs.AI 84%

\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding

Bin Lei, Nuo Xu, Ali Payani, Mingyi Hong, Chunhua Liao, Yu Cao, Caiwen Ding

机构 * University of Minnesota(明尼苏达大学) Cisco Research(思科研究) Lawrence Livermore National Labs(劳伦斯利弗莫尔国家实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11368 2025-10-07 cs.CV 83%

From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation

Jingkun Chen, Haoran Duan, Xiao Zhang, Boyan Gao, Vicente Grau, Jungong Han

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学) Department of Automation, Tsinghua University(自动化系,清华大学) School of Information Science and Technology, Northwest University(信息科学与技术学院,西北大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03840 2025-10-07 cs.CV 79%

Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models

Pranav Sharma, Shivank Garg, Durga Toshniwal

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥里分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments ACM MM'25, MALLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03376 2025-10-07 cs.CV eess.IV 79%

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

Sanjukta Ghosh

专题命中 视觉定位与Grounding :visual language model(title,abstract);分类 cs.CV

Comments Pre-review version submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 76%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23258 2025-10-07 cs.CV 74%

OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting

Atakan Topaloglu, Kunyi Li, Michael Niemeyer, Nassir Navab, A. Murat Tekalp, Federico Tombari

机构 * ETH Zürich(苏黎世联邦理工学院) Koç University(科卡大学) KUIS AI Center(KUIS人工智能中心) Technical University of Munich(慕尼黑技术大学) Google(谷歌) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Project page available at: https://atakan-topaloglu.github.io/oraclegs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17931 2025-10-07 cs.CV cs.AI cs.LG 67%

AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models

Xingjian Li, Qifeng Wu, Adithya S. Ubaradka, Yiran Ding, Colleen Que, Runmin Jiang, Jianhua Xing, Tianyang Wang, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Brown University(布朗大学) National Institute of Technology Karnataka(印度卡纳塔克国家理工学院) University of Pittsburgh(匹兹堡大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04819 2025-10-07 cs.CV cs.CL 57%

Visual Representations inside the Language Model

Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin, Winson Han, Ranjay Krishna

机构 * University of Washington(华盛顿大学) University of California Los Angeles(加州大学洛杉矶分校) Allen Institute for AI(人工智能研究院)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04623 2025-10-07 cs.AI 57%

MedPAO: A Protocol-Driven Agent for Structuring Medical Reports

Shrish Shrinath Vaidya, Gowthamaan Palani, Sidharth Ramesh, Velmurugan Balasubramanian, Minmini Selvam, Gokulraja Srinivasaraja, Ganapathy Krishnamurthi

机构 * Department of Data Science and AI, IIT Madras, India(数据科学与人工智能系,印度理工学院马德拉斯学院) Department of Engineering Design, IIT Madras, India(工程设计系,印度理工学院马德拉斯学院) LoveForm Health Technologies, India(LoveForm健康科技公司,印度) Department of Radiology and Imaging Sciences, Sri Ramachandra Institute of Higher Education and Research, India(放射学与成像科学系, Sri Ramachandra高等教育与研究学院,印度) Department of Neuro and Interventional Radiology, Sri Ramachandra Institute of Higher Education and Research, India(神经放射学与介入放射学系,Sri Ramachandra高等教育与研究学院,印度)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Paper published at "Agentic AI for Medicine" Workshop, MICCAI 2025

Journal ref Lecture Notes in Computer Science, vol 16147, 2025. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03955 2025-10-07 cs.CV 57%

Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs

Sameep Vani, Shreyas Jena, Maitreya Patel, Chitta Baral, Somak Aditya, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23054 2025-10-07 cs.CV 57%

Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis

Ruilang Wang, Shuotong Xu, Bowen Liu, Runlin Huang, Donglong Chen, Weifeng Su

机构 * Beijing Normal–Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18710 2025-10-07 cs.AI 57%

Autonomous Data Agents: A New Opportunity for Smart Data

Yanjie Fu, Dongjie Wang, Wangyang Ying, Xinyuan Wang, Xiangliang Zhang, Huan Liu, Jian Pei

机构 * Arizona State University(亚利桑那州立大学) University of Kansas(堪萨斯大学) University of Notre Dame(圣母大学) Duke University(杜克大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18569 2025-10-07 cs.CV 57%

VisualChef: Generating Visual Aids in Cooking via Mask Inpainting

Oleh Kuzyk, Zuoyue Li, Marc Pollefeys, Xi Wang

机构 * ETH Zürich(苏黎世联邦理工学院) Microsoft(微软) TU Munich(慕尼黑工业大学) MCML

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments GCPR 2025 (oral presentation; Best Master's Thesis Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21815 2025-10-07 cs.IR cs.AI cs.CL 57%

Scientific Paper Retrieval with LLM-Guided Semantic-Based Ranking

Yunyi Zhang, Ruozhen Yang, Siqi Jiao, SeongKu Kang, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Korea University(韩国大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03455 2025-10-07 cs.CV 57%

PEaRL: Pathway-Enhanced Representation Learning for Gene and Pathway Expression Prediction from Histology

Sejuti Majumder, Saarthak Kapse, Moinak Bhattacharya, Xuan Xu, Alisa Yurovsky, Prateek Prasanna

机构 * Department of Biomedical Informatics, Stony Brook University, NY, USA(生物医学信息学系,石溪大学,纽约,美国) Department of Computer Science, Stony Brook University, NY, USA(计算机科学系,石溪大学,纽约,美国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03294 2025-10-07 cs.CV 57%

Domain-Robust Marine Plastic Detection Using Vision Models

Saanvi Kataria

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 16 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04569 2025-10-07 q-fin.TR 50%

Risk-Sensitive Option Market Making with Arbitrage-Free eSSVI Surfaces: A Constrained RL and Stochastic Control Bridge

Jian'an Zhang

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 34 pages including appendices; figures included. Primary subject class: q-fin.TR. Cross-lists: cs.LG; q-fin.CP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04049 2025-10-07 cs.PL 50%

Encoding Numeric Computations and Infusing Heuristic Knowledge Using Integrity Constraints in stableKanren

Xiangyu Guo, Ajay Bansal

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 12 pages, 2 figures, ICFP '25 The miniKanren and Relational Programming Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04041 2025-10-07 cs.RO 50%

SITCOM: Scaling Inference-Time COMpute for VLAs

Ayudh Saxena, Harsh Shah, Sandeep Routray, Rishi Rajesh Shah, Esha Pahwa

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI (SpaVLE). *Equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05135 2025-10-07 cs.RO 50%

LERa: Replanning with Visual Feedback in Instruction Following

Svyatoslav Pchelintsev, Maxim Patratskiy, Anatoly Onishchenko, Alexandr Korchemnyi, Aleksandr Medvedev, Uliana Vinogradova, Ilya Galuzinsky, Aleksey Postnikov, Alexey K. Kovalev, Aleksandr I. Panov

机构 * MIPT(莫斯科国立交通大学) Sberbank of Russia, Robotics Center(俄罗斯储蓄银行机器人中心) AIRI

专题命中 视觉定位与Grounding :visual language model(abstract)

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏