arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-06 至 2025-11-06 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2505.16495 2025-11-06 cs.CV 70%

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

Lingfeng Wang, Hualing Lin, Senda Chen, Tao Wang, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03549 2025-11-06 cs.SE cs.AI 57%

Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding

Ziv Nevo, Orna Raz, Karen Yorav

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 7 pages, 6 figures, to be published in AISM 2025, see https://aism25.github.io/aism25/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04789 2025-11-06 cs.CV 57%

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Microsoft(微软)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11519 2025-11-06 cs.CV cs.CL 57%

Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models

Hao Cheng, Erjia Xiao, Yichi Wang, Lingfeng Zhang, Qiang Zhang, Jiahang Cao, Kaidi Xu, Mengshu Sun, Xiaoshuai Hao, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) The Hong Kong University of Science and Technology(香港科学与技术大学) Beijing University of Technology(北京工业大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments This paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03186 2025-11-06 cs.AI 57%

Adobe Summit Concierge Evaluation with Human in the Loop

Yiru Chen, Sally Fang, Sai Sree Harsha, Dan Luo, Vaishnavi Muppala, Fei Wu, Shun Jiang, Kun Qian, Yunyao Li

机构 * Adobe Inc.(Adobe公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by 6th Workshop on Data Science with Human in the Loop @ VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03181 2025-11-06 cs.RO cs.LG 57%

Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control

Rewida Ali, Cristian C. Beltran-Hernandez, Weiwei Wan, Kensuke Harada

机构 * Department of Systems Innovation, Graduate School of Engineering Science, Osaka University(大阪大学系统创新部门,工学研究科) OMRON SINIC X Corporation(OMRON SINIC X公司) The National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03047 2025-11-06 cs.LG 57%

Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions

Emi Soroka, Tanmay Chopra, Krish Desai, Sanjay Lall

机构 * Department of Electrical Engineering Stanford University(电气工程系 斯坦福大学) Emissary Technologies

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07661 2025-11-06 cs.CV cs.RO 57%

ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones

Anurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong, Juan Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03165 2025-11-06 cs.RO 50%

SENT Map -- Semantically Enhanced Topological Maps with Foundation Models

Raj Surya Rajendran Kathirvel, Zach A Chavis, Stephen J. Guy, Karthik Desingh

机构 * Minnesota Robotics Institute (MnRI)(明尼苏达州机器人研究所) Department of Computer Science and Engineering (CS&E)(计算机科学与工程系) University of Minnesota(明尼苏达大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏