发表机构
School of Computer Science, Northwestern Polytechnical University; The Hong Kong Polytechnic University; City University of Hong Kong; Department of Electrical and Computer Engineering, National University of Singapore; Generative AI Research and Development Center, The Hong Kong University of Science and Technology(西北工业大学计算机学院; 香港理工大学; 香港城市大学; 新加坡国立大学电气与计算机工程系; 香港科技大学生成式人工智能研发中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有世界模型评估忽略灾难性故障风险的问题,提出 BasinLens 方法,可在有限查询预算下发现自然输入引发的可复现、局部持续的世界模型故障,暴露常规基准未覆盖的漏洞。
AI 中文摘要
世界模型可预测与动作相关的未来状态,是下游规划与控制任务的关键内部模拟器。然而,世界模型的灾难性预测故障会通过控制流程危险地传播,因为后续智能体或模型的训练与决策高度依赖这些世界模型预测的环境演化。现有评估忽略了这种系统性风险:通过汇总一般查询下良性生成内容的平均误差,它们未能在罕见或未观测到的条件-动作组合下对模型进行灾难性崩溃的压力测试。为填补这一空白,我们将自然输入故障发现问题形式化:在有限查询预算下,寻找能引发严重预测风险的环境有效条件与动作前缀,验证这些故障是否在新随机种子上复现,并测试其在附近有效编辑下的持续性。发现此类关键故障在计算上极具挑战性,因为有效条件-动作组合会呈指数级增长,在带噪声的 rollout 成本高昂的情况下,穷尽搜索或标准采样均不可行。为解决这一问题,我们提出 BasinLens,该方法利用有效输入的底层结构——每个坐标具有环境定义的语义类型与允许域,将不确定性引导的全局搜索与类型化局部替换相结合。在多种基准测试与世界模型族中,BasinLens 揭示了常规评估未能发现的可复现且局部持续的故障模式,表明平均基准测试会掩盖世界模型驱动控制中的重要漏洞。
英文摘要
World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic prediction failures of world models could dangerously propagate through the control pipeline, as subsequent agent or model training and decision-making depend heavily on the continuous environment evolution forecasted by these world models. Existing evaluations overlook this systemic risk: by aggregating average errors over benign generations from general queries, they fail to stress-test the model against catastrophic collapses under rare or unobserved condition-action combinations. To bridge this gap, we formalize the natural-input failure discovery problem: under a finite query budget, finding environment-valid conditions and action prefixes that induce severe prediction risk, verifying whether these failures reproduce on fresh seeds, and testing their persistence under nearby valid edits. Discovering such critical failures is computationally challenging, as valid condition-action combinations explode exponentially, rendering exhaustive search or standard sampling infeasible given the high cost of noisy rollouts. To tackle this, we propose BasinLens, which exploits the underlying structure of valid inputs, where each coordinate possesses environment-defined semantic types and admissible domains, by pairing uncertainty-guided global search with typed local replacements. Across diverse benchmarks and world-model families, BasinLens exposes reproducible and locally persistent failure modes that conventional evaluations fail to reveal, showing that average-case benchmarks can mask important vulnerabilities in world-model-driven control.