arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-22 至 2026-01-22 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 2 篇

2601.14732 2026-01-22 cs.CV cs.CL cs.MM 57%

DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling

DeepMoLM: 利用视觉和几何结构信息进行分子-文本建模

Jing Lan, Hexiao Ding, Hongzhao Chen, Yufeng Jiang, Nga-Chun Ng, Gwing Kei Yip, Gerald W. Y. Cheng, Yunlin Mao, Jing Cai, Liang-ting Lin, Jung Sun Yoo

机构 * Department of Health Technology and Informatics, The Hong Kong Polytechnic University(健康科技与信息学系,香港理工大学) Department of Nuclear Medicine and PET, Hong Kong Sanatorium and Hospital(核医学与PET部,香港疗养院及医院) Department of Diagnostic and Interventional Radiology, Queen Elizabeth Hospital Hong Kong SAR, China(诊断与介入放射学部,香港特别行政区中国女王医院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 DeepMoLM通过双视角框架结合视觉和几何信息,提升分子-文本建模的准确性与物理合理性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14874 2026-01-22 cs.RO 50%

HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation

HumanoidVLM: 用于接触密集型人形机器人工操作的视觉-语言引导阻抗控制

Yara Mahmoud, Yasheerah Yaqoot, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

机构 * Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所)

专题命中 其他VLM :vision-language model(abstract)

AI总结 HumanoidVLM通过视觉-语言模型和检索增强生成模块,实现人形机器人在接触密集场景中的自适应阻抗控制与抓取配置选择。

Comments This paper has been accepted for publication at LBR of HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏