MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
MIRAGE: 具有隐式推理和生成世界模型的移动智能体
Zhichao Yang, Yuanze Hu, Haojie Hao, Longkun Hao, Dongshuo Huang, Hongyu Lin, Gen Li, Lanqing Hong, Yihang Lou, Yan Bai
机构
*
Beihang University(北京航空航天大学)
;
Northwestern Polytechnical University(西北工业大学)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
Independent Researcher(独立研究者)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.AI
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
WorldMemArena: 通过动作-世界交互评估多模态智能体记忆
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu, Yepeng Liu, Lin Long, Yichen Guo, Nuo Chen, Zhaotian Weng, Elena Kochkina, Simerjot Kaur, Charese Smiley, Xiaomo Liu, James Zou, Sheng Liu, Yuheng Bu, Songyou Peng, Xin Eric Wang
机构
*
University of California, Santa Barbara(加州大学圣芭芭拉分校)
;
J.P. Morgan Chase(摩根大通)
;
ETH Zurich(苏黎世联邦理工学院)
;
Stanford University(斯坦福大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Carnegie Mellon University(卡内基梅隆大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.CV
AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing
AIGaitor: 面向所有人的隐私保护与无云端运动分析——基于边缘计算
Lauhitya Reddy, Trisha M. Kesar, Hyeokhyen Kwon
机构
*
Department of Biomedical Informatics, Emory University(埃默里大学生物医学信息学系)
;
Department of Rehabilitation Medicine, Emory University(埃默里大学康复医学系)
;
The Wallace H. Coulter Department of Biomedical Engineering, Emory University and Georgia Institute of Technology(埃默里大学和佐治亚理工学院的Wallace H. Coulter生物医学工程系)
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
面向具身操作的高效视觉-语言-动作模型:系统综述
Weifan Guan, Qinghao Hu, Aosheng Li, Jian Cheng
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
AiRiA
;
Nanjing University of Information Science and Technology(南京信息科学技术大学)
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing
VLA-Trace: 通过表示与行为追踪诊断视觉-语言-动作模型
Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju
机构
*
University of Science and Technology of China(中国科学技术大学)
;
University of Manchester(曼彻斯特大学)
;
Beihang University(北航)
;
Fudan University(复旦大学)
;
University of New South Wales(新南威尔士大学)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
面向视觉-语言导航的不确定性感知高斯地图
Jianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao, Yingxue Zhang, Zhanguang Zhang, Sida Peng, Yi Yang, Wenguan Wang
机构
*
The State Key Lab of Brain-Machine Intelligence(脑机智能国家重点实验室)
;
Department of Foundation model, 2012 Labs, Huawei(基础模型部门,2012实验室,华为)
;
Noah’s Ark Lab, 2012 Labs, Huawei(诺亚方舟实验室,2012实验室,华为)
;
School of Software Technology, Zhejiang University(浙江大学软件学院)
IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
IntentionNav: 一种基于隐式人类指令的意图驱动目标导航基准
Lin Qian, Shijie Li, Sihao Lin, Xuan Zhang, Bangya Liu, Yanran Li, Hujun Yin
机构
*
The University of Manchester(曼彻斯特大学)
;
A*STAR
;
Responsible AI Research Centre, Adelaide University(阿德莱德大学负责任人工智能研究中心)
;
University of Bedfordshire(贝福德郡大学)