arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-23 至 2025-10-23 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 2 篇

2510.19307 2025-10-23 cs.CV 79%

Unified Reinforcement and Imitation Learning for Vision-Language Models

Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu

机构 * NVIDIA KAIST(韩国科学技术院) National Taiwan University(国立台湾大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments NeurIPS 2025, Project page: https://byungkwanlee.github.io/RIL-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17150 2025-10-23 cs.RO 67%

OmniVIC: A Self-Improving Variable Impedance Controller with Vision-Language In-Context Learning for Safe Robotic Manipulation

Heng Zhang, Wei-Hsing Huang, Gokhan Solak, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genoa, Italy(人机交互实验室,意大利理工学院,热那亚) Georgia Institute of Technology, Atlanta, USA(佐治亚理工学院,美国亚特兰大)

专题命中 幻觉与鲁棒性 :vision language model(abstract);VLM(abstract)

Comments Code, video and RAG dataset are available at \url{https://sites.google.com/view/omni-vic}

详情

展开后加载摘要…

URL PDF HTML 收藏