MEND:世界模型中潜在幻觉的无标签检测、定位与纠正
MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
浏览论文内容
中文总结 AI 辅助
针对世界模型中的潜在幻觉,提出MEND方法,利用单个分数网络同时实现无标签检测、定位与纠正,在导航任务中AUROC达0.80。
中文摘要 AI 辅助
世界模型正成为计算机视觉的下一个主要前沿方向。然而,其鲁棒性目前在很大程度上尚未被探索。我们识别了潜在世界模型中的幻觉现象:给定一个状态和一个动作,预测的下一个潜在向量可能解码为一个从未发生过的场景。由于该预测在统计上表现正常,并且被模型自回归地反馈,因此该错误既是无声的,又是累积的。我们研究在推理时,针对一个冻结的自监督世界模型,在缺乏真实错误标签的情况下,这种潜在幻觉是否可以被检测、定位和纠正。我们引入了掩蔽经验贝叶斯神经去噪(MEND),这是一个通过去噪分数匹配在真实转移上训练的条件分数网络,其分数场扮演三个角色:其幅度检测幻觉,其逐令牌字段将幻觉定位到特定图像块,并定义了推理时的纠正方向。在两个导航环境中,MEND在不使用动作的情况下,以高达0.80的AUROC检测幻觉,超过了单高斯密度基线,同时还能定位错误(逐令牌AUPRC高达0.87)并纠正错误,所有这些都来自一个分数场。我们的纠正可靠地减少了单步潜在错误并改善了预测。我们发现部分错误与数据流形相切,因此我们专注于检测和定位,同时强调纠正的前景。
英文摘要
World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
发表机构
- University of Melbourne(墨尔本大学)
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。