CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
CodeGraphVLP:代码规划器与语义图状态的结合用于非马尔可夫视觉-语言-动作模型
Khoa Vo, Sieu Tran, Taisei Hanyu, Yuki Ikebe, Duy Nguyen, Nghi D. Q. Bui, Minh Vu, Anthony Gunderman, Chase Rainwater, Anh Nguyen, Ngan Le
机构
*
University of Arkansas(亚拉巴马大学)
;
Max Planck Research School for Intelligent Systems and the University of Stuttgart(马克斯·普朗克智能系统研究学校和斯图加特大学)
;
Center of AI Research, VinUniversity(Vin大学人工智能研究中心)
;
TU Wien(维也纳技术大学)
;
University of Liverpool(利物浦大学)
SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis
SwipeGen: 通过类人滑动合成弥合GUI代理的执行差距
Xuan Wang, Siyuan Su, Quantong Fu, Yongxiang Hu, Yangfan Zhou
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室)
机构
*
Sichuan University(四川大学)
;
Sun Yat-sen University(中山大学)
;
Australian National University(澳大利亚国立大学)
;
Peking University(北京大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Shanghai Jiao Tong University(上海交通大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Beihang University(北京航空航天大学)
;
Huawei Foundation Model Department(华为基础模型部门)
MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
MAC 2026:推动微动作分析迈向细粒度理解
Kun Li, Dan Guo, Jihao Gu, Pengyu Liu, Xiaobai Li, Haoyu Chen, Yanbin Hao, Guoying Zhao, Meng Wang
机构
*
United Arab Emirates University(阿联酋大学)
;
Hefei University of Technology(合肥工业大学)
;
University College London(伦敦大学学院)
;
Zhejiang University(浙江大学)
;
University of Oulu(奥卢大学)
;
CMVS, University of Oulu(奥卢大学计算机视觉与媒体研究中心)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.CV
CommentsWithdrawn due to ongoing technical improvements. The work requires further refinement and additional experiments to meet our quality standards. A revised version will be submitted in the future
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
Monash University(莫纳什大学)
;
Pusan National University(釜山国立大学)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
Comments8 pages main text, 21 pages total including appendices; 11 figures, 7 tables, 2 algorithms. Benchmark, harness, and model checkpoints to be released