arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2508.17724 2025-08-26 cond-mat.soft 50%

Adhesion Control through Electric Field-Induced Water Adsorption at Oxidized Silicon Interfaces

Tunç Çiftçi, Jonathon Cottom, Rachid Hahury, Emilia Olsson, Bart Weber

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05718 2025-08-26 cs.CL 50%

A Factuality and Diversity Reconciled Decoding Method for Knowledge-Grounded Dialogue Generation

Chenxu Yang, Zheng Lin, Chong Tian, Liang Pang, Lanrui Wang, Zhengyang Tong, Qirong Ho, Yanan Cao, Weiping Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09071 2025-08-25 cs.CL 50%

Exploration of Plan-Guided Summarization for Narrative Texts: the Case of Small Language Models

Matt Grenander, Siddharth Varia, Paula Czarnowska, Yogarshi Vyas, Kishaloy Halder, Bonan Min

机构 * AWS AI Labs(AWS人工智能实验室) School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to the 7th Workshop on Narrative Understanding (WNU), co-located with NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14056 2025-08-21 cs.CL cs.DB 50%

Confidence Estimation for Text-to-SQL in Large Language Models

Sepideh Entezari Maleki, Mohammadreza Pourreza, Davood Rafiei

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12388 2025-08-19 cs.HC 50%

When motivation can be more than a message: designing agents to boost physical activity

Alessandro Silacci, Maurizio Caon, Mauro Cherubini

专题命中 视觉定位与Grounding :grounding(abstract)

Comments This is a pre-peer-review version of a paper with the same title accepted at 20th IFIP TC13 International Conference on Human-Computer Interaction (INTERACT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09095 2025-08-19 cond-mat.stat-mech 50%

Benchmarking Energy Calculations Using Formal Proofs

Ejike D. Ugwuanyi, Colin T. Jones, John Velkey, Tyler R. Josephson

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Molecular Physics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10444 2025-08-15 cs.CL 50%

DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales

Herun Wan, Jiaying Wu, Minnan Luo, Xiangzheng Kong, Zihan Ma, Zhi Zeng

专题命中 视觉定位与Grounding :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08737 2025-08-13 cs.HC 50%

From Data to Insight: Using Contextual Scenarios to Teach Critical Thinking in Data Visualisation

Jonathan C. Roberts, Peter Butcher, Panagiotis D. Ritsos

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 6 pages, 5 figures, IEEE VIS Workshop on Visualization Education, Literacy, and Activities 2025, Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21110 2025-08-13 gr-qc astro-ph.CO hep-th 50%

The frozen vanilla model: Exploring dark sector interactions with delta effective theories

Martín G. Richarte, Luiz F. Guimarães, Susana J. Landau, Júlio C. Fabris

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 37 pages (double column format), 18 figures, and 4 appendices. References updated and version improved; accepted for publication in PRD

Journal ref Phys. Rev. D 112, 043510 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08084 2025-08-12 cs.CY 50%

$100,000 or the Robot Gets it! Tech Workers' Resistance Guide: Tech Worker Actions, History, Risks, Impacts, and the Case for a Radical Flank

Mohamed Abdalla

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07786 2025-08-12 math.LO cs.LO 50%

Proof-theoretic Semantics for Second-order Logic

Alexander V. Gheorghiu, David J. Pym

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05534 2025-08-08 cs.CL 50%

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim, Pranav Shetty, Dongsheng Wang, Zhiqiang Ma, Armineh Nourbakhsh, Xiaomo Liu

机构 * School of Computation, Information, and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学) Graduate Institute of International and Development Studies, Geneva(国际与发展研究研究生院,日内瓦) JPMorgan AI Research(摩根大通人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to ACL 2025-Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05061 2025-08-08 cs.DB cs.IR 50%

Data-Aware Socratic Query Refinement in Database Systems

Ruiyuan Zhang, Chrysanthi Kosyfaki, Xiaofang Zhou

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02068 2025-08-08 cs.RO 50%

"Set It Up": Functional Object Arrangement with Compositional Generative Models (Journal Version)

Yiqing Xu, Jiayuan Mao, Linfeng Li, Yilun Du, Tomas Lozáno-Pérez, Leslie Pack Kaelbling, David Hsu

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments This is the journal version accepted to the International Journal of Robotics Research (IJRR). It extends our prior work presented at Robotics: Science and Systems (RSS) 2024, with a new compositional program induction pipeline from natural language, and expanded evaluations on personalized bookshelf and bedroom furniture layout tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05236 2025-08-08 cs.MA 50%

Towards Language-Augmented Multi-Agent Deep Reinforcement Learning

Maxime Toquebiau, Jae-Yun Jun, Faïz Benamar, Nicolas Bredeche

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accespted at the European Conference on Artificial Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00099 2025-08-07 cs.CY cs.MA physics.soc-ph 50%

Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior

Egil Diau

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Position paper on LLM-agent simulation of financial structures. This update clarifies setup and adds a reciprocity-based table. Builds on arXiv:2505.02945 and 2505.08319

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04598 2025-08-07 cs.RO 50%

$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything

Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Xinyu Zheng, Pengwei Wang, Zhongyuan Wang, Wenbo Ding, Shanghang Zhang

专题命中 视觉定位与Grounding :VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04504 2025-08-07 cs.CY 50%

Moving beyond harm. A critical review of how NLP research approaches discrimination

Katrin Schulz, Marjolein Lanzing, Giulia Martinez Brenner

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04183 2025-08-07 cs.CL 50%

Characterizing Deep Research: A Benchmark and Formal Definition

Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi, Aaron Halfaker, Amit Deshpande, Navin Goyal, Ankur Gupta, Nagarajan Natarajan, Amit Sharma

机构 * Microsoft Research(微软研究院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments First three authors contributed equally (ordered alphabetically)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03830 2025-08-07 cs.PL 50%

If-T: A Benchmark for Type Narrowing

Hanwen Guo, Ben Greenman

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 2, Article 17

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03340 2025-08-06 cs.SE 50%

Key-Augmented Neural Triggers for Knowledge Sharing

Alex Wolf, Marco Edoardo Palma, Pooja Rani, Harald C. Gall

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07234 2025-08-05 math.AG 50%

Spectral Fingerprints of Algebraic Cycles: A Hodge-Theoretic Approach to the Hodge Conjecture and Special L-Values

Bita Hajebi, Pooya Hajebi

专题命中 视觉定位与Grounding :grounding(abstract)

Comments arXiv admin note: This submission has been withdrawn due to violation of arXiv policies for acceptable submissions

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20306 2025-07-29 math.NA cs.NA 50%

A Hybrid Particle-Continuum Method for Simulating Fast Ice via Subgrid Iceberg Interaction

Carolin Mehlmann, Saskia Kahl

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12370 2025-07-29 cs.CL 50%

Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs

Rupak Sarkar, Neha Srikanth, Taylor Hudson, Rachel Rudinger, Claire Bonial, Philip Resnik

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19493 2025-07-29 cs.HC eess.IV 50%

From Bench to Bedside: A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

Yaowei Bai, Ruiheng Zhang, Yu Lei, Jingfeng Yao, Shuguang Ju, Chaoyang Wang, Wei Yao, Yiwan Guo, Guilin Zhang, Chao Wan, Qian Yuan, Xuhua Duan, Xinggang Wang, Tao Sun, Yongchao Xu, Chuansheng Zheng, Huangxuan Zhao, Bo Du

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17227 2025-07-24 cond-mat.mes-hall physics.atm-clus 50%

Topological Zero Modes in Non-Hermitian Topolectrical Systems: Size and Impedance Control

S M Rafi-Ul-Islam, Zhuo Bin Siu, Md. Saddam Hossain Razo, Mansoor B. A. Jalil

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15244 2025-07-22 cs.HC 50%

How Does Empirical Research Facilitate Creation Tool Design? A Data Video Perspective

Leixian Shen, Leni Yang, Haotian Li, Yun Wang, Yuyu Luo, Huamin Qu

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00651 2025-07-22 cs.CY 50%

Innovative Tangible Interactive Games for Enhancing Artificial Intelligence Knowledge and Literacy in Elementary Education: A Pedagogical Framework

Nikolaos Sampanis

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12621 2025-07-21 cs.HC cs.GR cs.MA 50%

NLI4VolVis: Natural Language Interaction for Volume Visualization via LLM Multi-Agents and Editable 3D Gaussian Splatting

Kuangshi Ai, Kaiyuan Tang, Chaoli Wang

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments IEEE VIS 2025. Project Page: https://nli4volvis.github.io/

Journal ref IEEE Transactions on Visualization and Computer Graphics (TVCG), vol. 32, no. 1, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11906 2025-07-17 cs.MA cs.HC 50%

CoCre-Sam (Kokkuri-san): Modeling Ouija Board as Collective Langevin Dynamics Sampling from Fused Language Models

Tadahiro Taniguchi, Masatoshi Nagano, Haruumi Omoto, Yoshiki Hayashi

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏