AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
AERMANI-VLM:基于视觉语言模型的空中 manipulation 的结构提示与推理
Sarthak Mishra, Rishabh Dev Yadav, Avirup Das, Saksham Gupta, Wei Pan, Spandan Roy
机构
*
Robotics Research Center, IIIT Hyderabad(IIIT海得拉巴机器人研究中心)
;
Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系)
;
Newcastle University(纽卡斯尔大学)
专题命中
视觉推理
:VLM(title,abstract);vision language model(title)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
ThinkJEPA:赋予潜在世界模型大型视觉-语言推理能力
Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu
机构
*
Northeastern University(东北大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
University of Maryland(马里兰大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of Washington(华盛顿大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
ByteDance Inc.(字节跳动公司)
;
Zhejiang University of Technology(浙江工业大学)
;
Hong Kong Baptist University(香港浸会大学)
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
多模态焦点的潮起潮落:为基于视觉的语言模型推理调度视觉中继窗口
Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
机构
*
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Fashion and Textiles, The Hong Kong Polytechnic University(香港理工大学纺织及制衣学院)
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)