Simulating the Real World: A Unified Survey of Multimodal Generative Models
模拟现实世界:多模态生成模型的统一综述
Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, Ce Liu, Anyi Rao, Zeyu Wang, Hui Xiong
机构
*
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能前沿技术研究所,香港科学与技术大学(广州))
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(计算机科学与工程系,香港科学与技术大学香港特别行政区)
;
MMLab, The Hong Kong University of Science and Technology(多模态实验室,香港科学与技术大学)
;
School of Electronics and Communication Engineering, Shenzhen Campus of Sun Yat-sen University(电子与通信工程学院,中山大学深圳校区)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国)
;
Tsinghua University, Guangdong, China(清华大学,广东,中国)
;
Bosch (China) Investment Co., Ltd., Shanghai, China(博世(中国)投资有限公司,上海,中国)
DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
DriveFine: 通过细化增强的掩码扩散VLA实现精确且鲁棒的驾驶
Chenxu Dang, Sining Ang, Yongkang Li, Haochen Tian, Jie Wang, Guang Li, Hangjun Ye, Jie Ma, Long Chen, Yan Wang
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Xiaomi EV(小米电动车)
;
Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学)
机构
*
National Yang Ming Chiao Tung University(国立阳明交通大学)
;
National Kaohsiung Normal University(国立高雄师范大学)
;
Institute for Clarity in Documentation(文档清晰研究所)
;
Inria Paris-Rocquencourt(巴黎-勒克努尔研究所)
;
Rajiv Gandhi University(拉贾·甘地大学)
;
Tsinghua University(清华大学)
;
Palmer Research Laboratories(帕勒尔研究实验室)
AI总结
HyperRAG通过n元超图推理提升检索增强生成的准确性与效率
CommentsAccepted by The ACM Web Conference 2026 (WWW '26)
Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding
文本优先于视觉:针对超高清遥感理解的代理强化学习在超高清遥感理解中的知识注入至关重要
Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yuhao Zhou, Di Wang, Yifan Zhang, Haoyu Wang, Haiyan Zhao, Hongda Sun, Long Lan, Jun Song, Yulin Wang, Jing Zhang, Wenlong Zhang, Bo Du
机构
*
National University of Defense Technology, China(国防科技大学)
;
Beijing University of Posts and Telecommunications, China(北京邮电大学)
;
University of the Chinese Academy of Sciences, China(中国科学院大学)
;
Sichuan University, China(四川大学)
;
Wuhan University, China(武汉大学)
;
Chinese Academy of Science, China(中国科学院)
;
Tsinghua University, China(清华大学)
;
Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)
;
Renmin University of China, China(中国人民大学)
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
HiVid: 基于大语言模型的视频显著性识别用于内容感知点播与直播
Jiahui Chen, Bo Peng, Lianchen Jia, Zeyu Zhang, Tianchi Huang, Lifeng Sun
机构
*
Tsinghua University(清华大学)
;
The Australian National University(澳大利亚国立大学)
;
Key Laboratory of Pervasive Computing, Ministry of Education(教育部普适计算重点实验室)
Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind
超越词语:通过心灵理论评估和弥合用户代理交互中的认知分歧
Minyuan Ruan, Ziyue Wang, Kaiming Liu, Yunghwei Lai, Peng Li, Yang Liu
机构
*
Dept. of Comp. Sci. \& Tech., Institute for AI, Tsinghua University, Beijing, China
;
Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
R-Diverse: 缓解自博弈LLM训练中的多样性幻觉
Gengsheng Li, Jinghan He, Shijie Wang, Dan Zhang, Ruiqi Liu, Renrui Zhang, Zijun Yao, Junfeng Fang, Haiyun Guo, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Wuhan AI Research(武汉人工智能研究所)
;
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints
GRIP:通过几何路由约束实现混合专家的算法无关机器反学习
Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li
机构
*
School of Computer Science(计算机科学学院)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Department of Electrical Engineering(电气工程系)
;
Tsinghua University(清华大学)
;
School of Electrical and Computer Engineering(电气与计算机工程学院)
机构
*
School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
机构
*
College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine, Shanghai, China(上海交通大学医学院健康科学与技术学院)
;
Fudan University, Shanghai, China(复旦大学)
;
Shanghai Innovation Institute, Shanghai, China(上海创新研究院)
;
Tsinghua University, Beijing, China(清华大学)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)
;
Ruijin Hospital, Shanghai Jiaotong University, Shanghai, China(上海交通大学瑞金医院)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Zhongguancun Academy, Beijing, China(中关村学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Tsinghua University(清华大学)
;
Eastern Institute of Technology, Ningbo(宁波东部科技研究院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nankai University(南开大学)
Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
基于CLIP引导的多模态扩散模型的图像到脑信号生成用于视觉假体
Ganxi Xu, Zhao-Rong Lai, Yuting Tang, Yonghao Song, Guoxu Zhou, Boyu wang, Jian Zhu, Jinyi Long
机构
*
Jinan University, Guangzhou, China
;
Department of Rehabilitation Medicine, The First Affiliated Hospital of Jinan University, Guangzhou, China
;
Western University, Ontario, Canada
;
Guangdong University of Technology, Guangzhou, China
;
Tsinghua University, Beijing, China