机构
*
Sun Yat-sen University(中山大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Baidu Inc.(百度公司)
;
Cardiff University(卡迪夫大学)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
为什么我们需要世界模型来实现通用人工智能:大语言模型失败之处以及世界模型如何可能超越
Feisal Alaswad, Batoul Aljaddouh, Maher Alrahhal, Poovammal E, Talal Bonny
机构
*
Department of Computing Technologies(计算技术系)
;
SRM Institute of Science and Technology(SRM科学与技术学院)
;
Bio-Sensing and Bio-Sensors Group(生物传感与生物传感器组)
;
Smart Automation and Communication Technologies Research Institute of Sciences and Engineering(科学与工程智能自动化与通信技术研究所)
;
University of Sharjah, UAE(阿联酋沙迦大学)
;
Department of Computer Engineering(计算机工程系)
;
College of Computing and Informatics(计算与信息学院)
Self-supervised Hierarchical Visual Reasoning with World Model
基于世界模型的自监督分层视觉推理
Yuanfei Xu, Lin Liu, Wengang Zhou, Mingxiao Feng, Houqiang Li
机构
*
Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心)
机构
*
Tencent Youtu Lab(腾讯优图实验室)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
;
University of Warwick(沃林汉大学)
;
Monash University(墨尔本大学)
;
The Hong Kong Polytechnic University(香港理工大学)
SURGE: Approximation and Training Free Particle Filter for Diffusion Surrogate
SURGE: 扩散替代模型的近似与免训练粒子滤波
Lifu Wei, Yinuo Ren, Naichen Shi, Yiping Lu
机构
*
Department of Mechanical Engineering, Northwestern University, Evanston, IL, United States
;
Institute for Computational \& Mathematical Engineering, Stanford University, Stanford, CA, United States
;
Department of Industrial Engineering \& Management Sciences, Northwestern University, Evanston, IL, United States
Federated Semantic Knowledge Graphs for Laboratory Workflows: A Structured Expert Elicitation Methodology Demonstrated Through Bioanalytical Workflow Twins
面向实验室工作流的联邦语义知识图谱:通过生物分析工作流孪生展示的结构化专家启发方法
Luis F. Schachner, Vinith Thamizhazhagan, Sara Tanenbaum, John C. Tran, Pamela P. F. Chan, Mandy Kwong, Andy Chang, Maureen Beresini, Margaret Porter Scott
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
WorldCraft: 从相机导航到交互式视频世界模型中的物体操控
Bohai Gu, Taiyi Wu, Yueyang Yuan, Jian Liu, Xiaocheng Lu, Dazhao Du, Jie Zhang, Jinxiang Lai, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
AI Technology Center, Tencent Video, Tencent(腾讯视频AI技术中心,腾讯)
;
Wuhan University(武汉大学)
;
Peking University(北京大学)
专题命中
视频世界模型
:world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
TimeSpot: 在真实世界场景中评估视觉语言模型的地理时间理解能力
Azmine Toushik Wasi, Shahriyar Zaman Ridoy, Koushik Ahamed Tonmoy, Kinga Tshering, S. M. Muhtasimul Hasan, Wahid Faisal, Tasnim Mohiuddin, Md Rizwan Parvez
机构
*
Computational Intelligence and Operations Laboratory (CIOL), Bangladesh(计算智能与运筹实验室(CIOL),孟加拉国)
;
Shahjalal University of Science and Technology (SUST), Sylhet, Bangladesh(沙赫jalal科学与技术大学(SUST),沙赫里尔,孟加拉国)
;
North South University (NSU), Dhaka, Bangladesh(北南大学(NSU),达卡,孟加拉国)
;
Qatar Computing Research Institute (QCRI), Doha, Qatar(卡塔尔计算研究中心(QCRI),多哈,卡塔尔)