arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-21 至 2025-11-21 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 5 篇

2511.15720 2025-11-21 cs.AI 79%

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

利用大语言和视觉-语言模型进行施工工地自动危险检测

Islem Sahraoui

机构 * University of Houston Cullen College of Engineering Department of Civil and Environmental Engineering(德克萨斯大学休斯顿分校库伦工程学院土木与环境工程系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.AI

AI总结 本研究利用大语言和视觉-语言模型,通过分析文本和图像数据,提高施工工地的安全隐患检测效率。

Comments Master thesis, University of Houton

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10241 2025-11-21 cs.CV 79%

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

TubeRMC: 基于互约束的管状重建用于弱监督空间-时间视频定位

Jinxuan Li, Yi Zhang, Jian-Fang Hu, Chaolei Tan, Tianming Liang, Beihao Xia

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 TubeRMC通过引入互约束的管状重建方法,在弱监督条件下提升空间-时间视频定位的准确性和一致性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08546 2025-11-21 cs.RO 78%

Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback

通过闭环状态反馈接地LLMs用于机器人任务规划

Vineet Bhat, Ali Umut Kaypak, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 本文提出BrainBody-LLM方法,通过闭环反馈机制提升LLMs在机器人任务规划中的鲁棒性,实现在复杂任务中的显著性能提升。

Comments Preprint version. Accepted full paper available here: https://advanced.onlinelibrary.wiley.com/doi/10.1002/adrr.202500072

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 77%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00091 2025-11-21 cs.CY cs.AI 57%

A First-Principles Based Risk Assessment Framework and the IEEE P3396 Standard

基于第一性原理的风险评估框架及IEEE P3396标准

Richard J. Tong, Marina Cortês, Jeanine A. DeFalco, Mark Underwood, Janusz Zalewski

机构 * Chair, IEEE Artificial Intelligence Standards Committee (AISC) Institute of Astrophysics Space Sciences, University of Lisbon, Portugal Vice Chair, IEEE Artificial Intelligence Standards Committee University of New Haven Florida Gulf Coast University, United States State Academy of Applied Sciences, Ciechanow, Poland

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出基于第一性原理的生成式AI风险评估框架,通过信息分类系统识别不同风险类型并归责于相关方,旨在提升AI治理的严谨性与责任性。

Comments 8 pages with 3 tables. This manuscript is prepared for publication by the Institute of Electrical and Electronics Engineers, Standards Association (IEEE-SA), Sponsor Committee - Artificial Intelligence Standards Committee (C/AISC) as a White Paper of Working Group p3396 at https://standards.ieee.org/ieee/3396/11379/

Journal ref 2025 IEEE Conference on Artificial Intelligence (CAI), pp. 1588-1595, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏