发表机构
Carnegie Mellon University; Georgia Institute of Technology; Northeastern University; Adobe Research(卡内基梅隆大学; 佐治亚理工学院; 东北大学; 奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可审计的物理验证系统VeriPhy,通过纯文本规划器编译物理义务,结合冻结低级专家与三值状态裁决,在1500片段语料库评估中缺陷识别效果优于对比方法,可用于世界模型评估优化。
AI 中文摘要
生成视频的视觉流畅性并不意味着物理可靠性,单一的质量评分无法指明视频片段违反的义务或失效的时刻。本文提出VeriPhy,这是一种可审计的物理验证系统,其中纯文本规划器在观察任何帧之前,会将提示编译为类型化的物理义务和经静态验证的执行计划。执行期间,观察结果仅对声明的对冻结低级专家的调用进行门控和范围限定,这些专家包括分割与跟踪、计数、对所得轨迹的11种类型物理测量、深度、OCR以及音频事件检测。每个操作都会返回带有溯源信息的证据记录,其有效载荷在可用时为类型化测量值或显式标记的学习状态。类型化解析器和固定组合将可用记录映射为三值状态(支持、矛盾或未知,分别显示为合理、不合理或弃权(不执行)),并附带完整溯源,因此每个裁决都可追溯到产生它的证据。我们将评估锚定在包含1500个片段的语料库上,该语料库带有人类标注的缺陷记录,用于定位提示参考、空间和时间方面的真实生成失败。在包含304条此类记录的149个片段核心集上,VeriPhy识别出228条缺陷,而使用相同片段和相同声明的已发布问题分解评估器仅识别出164条。仅召回率无法将其与整体提示同一主干的方法区分开,后者达到222;将它们区分开的是,每个决策都保留其证据记录和溯源,使得每个裁决的轨迹都可审计,并可用作将批评裁决写回生成过程的接口。
英文摘要
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed. During execution, observations gate and scope only declared calls to frozen low-level experts (e.g., segmentation and tracking, counting, eleven typed physical measurements over the resulting tracks, depth, OCR, and audio-event detection). Each action returns a provenance-carrying evidence record whose payload, when usable, is either a typed measurement or an explicitly tagged learned state. Typed resolvers and fixed composition map usable records to a three-valued state (supported, contradicted, or unknown, surfaced as plausible, implausible, or abstain) with full provenance, so that every verdict is traceable to the evidence that produced it. We anchor evaluation in a 1,500-clip corpus of human-annotated flaw records that localize real generation failures in prompt reference, space, and time. On a 149-clip core carrying 304 such records, VeriPhy accounts for 228, against 164 for a published question-decomposition evaluator given the same clips and the same claims. Recall alone does not separate it from prompting the same backbone monolithically, which reaches 222; what separates them is that each decision retains its evidence record and provenance, making the traces auditable one verdict at a time and usable as the interface through which a critic verdict could be written back into generation.