arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-01 至 2025-12-01 共收录 15 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 15 篇

2511.23300 2025-12-01 cs.RO 85%

SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot

SafeHumanoid: 通过VLM-RAG驱动的人形机器人上半身阻抗控制

Yara Mahmoud, Jeffrin Sam, Nguyen Khang, Marcelino Fernando, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Haris Khan, Artem Lykov, Dzmitry Tsetserukou

机构 * Skolkovo Institute of Science and Technology(斯克洛尔沃科学与技术研究所)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision language model(abstract);grounding(abstract)

AI总结 SafeHumanoid通过VLM-RAG结合自身视觉,实现人形机器人上半身阻抗控制,提升人机交互的安全性与任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23311 2025-12-01 cs.CV cs.AI cs.CL 81%

Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach

迈向自动安全驾驶指令:一种大规模视觉语言模型方法

Haruki Sakajo, Hiroshi Takato, Hiroshi Tsutsui, Komei Soda, Hidetaka Kamigaito, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技术研究所) Teatis inc.(Teatis公司) Queensland university of technology(昆士兰理工大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种大规模视觉语言模型方法,用于生成安全驾驶指令,通过构建数据集并评估模型性能,展示了微调模型在自动驾驶安全中的应用与挑战。

Comments Accepted to MMLoSo 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23151 2025-12-01 cs.CV 79%

Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding

学习拒绝:用于视频时间定位中处理难相关查询的拒绝感知强化微调

Jin-Seop Lee, SungJoon Lee, SeongJun Jung, Boyang Li, Jee-Hyong Lee

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出RA-RFT方法,通过拒绝感知强化微调提升视频时间定位中对难相关查询的处理能力。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19221 2025-12-01 cs.CV cs.RO 77%

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

Percept-WAM:感知增强的环境感知-行动模型用于鲁棒的端到端自动驾驶

Jianhua Han, Meng Tian, Jiangtong Zhu, Fan He, Huixin Zhang, Sitong Guo, Dechang Zhu, Hao Tang, Pei Xu, Yuze Guo, Minzhe Niu, Haojie Zhu, Qichao Dong, Xuechao Yan, Siyuan Dong, Lu Hou, Qingqiu Huang, Xiaosong Jia, Hang Xu

机构 * Yinwang Intelligent Technology Co. Ltd.(亿网通智能科技有限公司) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 Percept-WAM通过整合2D/3D场景理解能力,提升自动驾驶的感知与行动决策,实现端到端鲁棒性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22256 2025-12-01 cs.CV 74%

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

UMind-VL:一种通用的超声视觉-语言模型,用于统一的 grounded perception 和全面的 interpretation

Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, Yuhang Liu, Dong Wang

机构 * Yizhun Medical AI Team(义诊医疗AI团队)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 UMind-VL 是一种通用超声视觉-语言模型,通过统一的 grounded perception 和 comprehensive interpretation 实现对医学影像的高效理解和诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22170 2025-12-01 cs.CV 70%

Partially Shared Concept Bottleneck Models

部分共享概念瓶颈模型

Delong Zhao, Qiang Huang, Di Yan, Yiqun Sun, Jun Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 PS-CBM通过部分共享的概念策略和概念效率准确性度量,提升模型的分类准确性和可解释性。

Comments 14 pages, 7 figures, 11 tables, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21984 2025-12-01 cs.CV 70%

PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation

PPBoost: 逐步提示增强用于文本驱动的医学图像分割

Xuchen Li, Hengrui Gu, Mohan Zhang, Qin Liu, Zhen Tan, Xinyuan Zhu, Huixue Zhou, Tianlong Chen, Kaixiong Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 PPBoost通过逐步增强弱文本提示为强空间指导,提升医学图像分割的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21705 2025-12-01 cs.CL cs.CV 70%

Insight-A: Attribution-aware for Multimodal Misinformation Detection

Insight-A: 多模态虚假信息检测中的归因意识

Junjie Wu, Yumeng Fu, Chen Gong, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 Insight-A通过归因意识和分层推理提升多模态虚假信息检测效果,有效识别伪造来源并增强跨模态一致性检查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09185 2025-12-01 cs.CV cs.AI 62%

A Neurosymbolic Framework for Interpretable Cognitive Attack Detection in Augmented Reality

面向增强现实的认知攻击检测的神经符号框架

Rongqian Chen, Allison Andreyev, Yanming Xiu, Joshua Chilukuri, Shunav Sen, Mahdi Imani, Bin Li, Maria Gorlatova, Gang Tan, Tian Lan

机构 * George Washington University(乔治华盛顿大学) Duke University(杜克大学) Pennsylvania State University(宾夕法尼亚州立大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CADAR框架,结合神经网络与符号推理,用于增强现实中的认知攻击检测,通过多模态表示和统计推理提升检测的可解释性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23146 2025-12-01 cs.CV 57%

InstanceV: Instance-Level Video Generation

InstanceV: 实例级视频生成

Yuheng Chen, Teng Hu, Jiangning Zhang, Zhucun Xue, Ran Yi, Lizhuang Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 InstanceV通过实例级控制和全局语义一致性,实现高质量视频生成,并在实例感知指标上超越现有最佳模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22906 2025-12-01 cs.CV 57%

See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection

见、排、滤:通过场景理解的重要词感知Clip过滤用于片段检索和亮点检测

YuEun Lee, Jung Uk Kim

机构 * YuEun Lee, Jung Uk Kim

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出通过识别查询中的重要词,结合多模态大语言模型实现视频片段检索和亮点检测的细粒度过滤方法,提升检索和检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22245 2025-12-01 cs.CV 57%

Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models

语义锚定用于文本到图像扩散模型中的稳健个性化

Seoyun Yang, Gihoon Kim, Taesup Kim

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 通过语义锚定方法,文本到图像扩散模型在有限参考图像中实现稳健个性化,平衡主题保真度与语义对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22025 2025-12-01 cs.CV 57%

Layover or Direct Flight: Rethinking Audio-Guided Image Segmentation

转折或直接飞行:重新思考音频引导的图像分割

Joel Alberto Santos, Zongwei Wu, Xavier Alameda-Pineda, Radu Timofte

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室、CAIDAS与IFI、乌尔姆大学、德国) Inria at Univ. Grenoble Alpes, CNRS, LJK, France(Inria于格勒诺布尔阿尔卑斯大学、CNRS、LJK、法国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出通过直接音频-视觉对齐进行图像分割,无需依赖文本转录,提升了鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21762 2025-12-01 cs.CL cs.AI 57%

Factors That Support Grounded Responses in LLM Conversations: A Rapid Review

支持LLM对话中基础响应的因素:一项快速回顾

Gabriele Cesar Iwashima, Claudia Susie Rodrigues, Claudio Dipolitto, Geraldo Xexéo

机构 * (June 2025)((2025年6月))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文通过快速回顾识别了支持LLM对话中基础响应的因素,重点分析了推理时间、训练后及强化学习方法,旨在提升LLM响应的准确性和可靠性。

Comments 28 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22426 2025-12-01 physics.optics 50%

Universal convolution from wave dynamics: photonic processing and encryption in synthetic dimension

通用卷积来自波动力学:在合成维度中光子处理与加密

Xiaolong Su, Weiwei Liu, Ruiqian Cheng, Haoru Zhang, Xinyao Guo, He Huang, Chengzhi Qin, Peixiang Lu, Bing Wang

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究通过光子合成晶格中的波动力学实现通用卷积,提出了一种基于卷积的光学加密方法,并展示了高吞吐量的图像处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏