SimVerity:当模拟智能体的成功能在物理部署中延续时?
SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
查看机构详情
- Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究量化模拟智能体成功与物理部署的关联,提出SimVerity框架,通过交叉验证判决转移,发现模拟通过与物理实际存在差异,可预测错误通过并提升智能体可审计性,暴露模拟器盲点。
中文摘要 AI 辅助
模拟评估被广泛用于对AI智能体进行基准测试,但模拟通过能在多大程度上为物理部署提供依据尚未得到系统量化。我们提出了SimVerity,这是一个判决转移保障框架:它在目标智能家居部署系统上重放匹配的场景,并对照独立合格的物理见证者对智能体执行情况进行交叉验证。我们的评估表明,部署成功是一个现实世界的过程,而非模拟中的静态属性:在同一执行过程中,完成情况、报告状态、可观测效果和最终结果存在差异。尽管某先进模拟器通过了全部240次灯光试验,但摄像头捕捉到了42次最终状态检查无法发现的亚秒级故障。错误通过是可预测的:从已测量试验中学习并在评估前锁定的风险剖面,能够预测其从未实际测量的路径上的故障,在两个队列的11次保留会话中均优于属性盲基线。智能体可审计性也可测量:切换某智能体循环的模型客户端/服务配置,将其场景匹配率从52-88%提升至100%。最后,第二个合格模拟器未提供独立交叉验证:它从未在任何重叠案例上产生分歧,只有物理测量暴露了它们共同的盲点。SimVerity将判决转移转化为明确决策:在部署前判定通过、弃权(不执行)或升级处理。
英文摘要
Simulated evaluation is widely used to benchmark AI agents, yet how much evidence a simulated pass provides about physical deployment has not been systematically quantified. We present SimVerity, a verdict-transfer assurance framework: it replays matched scenarios on target smart home deployments and cross-validates agent execution against independently qualified physical witnesses. Our evaluation highlights that deployment success is a real-world process, not a static property in simulation: completion, reported state, observable effect, and settled outcome diverged within the same execution. Although an advanced simulator cleared all 240 light trials, a camera caught 42 sub-second failures invisible to settled-state checks. False clearance was predictable: a risk profile learned from measured trials and locked before evaluation predicted failures on a path it never physically measured, beating a property-blind baseline in all eleven held-out sessions across two cohorts. Agent auditability was also measurable: switching one agent loop's model-client/serving configuration raised its scenario-matching share from 52-88% to 100%. Finally, a second qualified simulator added no independent cross-check: it never disagreed on any overlapping case, and only physical measurement exposed their shared blind spots. SimVerity turns verdict transfer into an explicit decision: clear, abstain, or escalate before deployment.