有界保真度模拟即演示阶段:用于治理基准的动捕交接
Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks
浏览论文内容
中文总结 AI 辅助
提出有界保真度模拟即演示阶段模式,通过抑制交接包络内接触物理,实现LLM驱动机器人治理基准的审计链字节级可重现,以近乎零成本提升基准可靠性。
中文摘要 AI 辅助
仿真到现实的研究将物理保真度作为首要目标:仿真器通过其再现真实接触动力学的程度来评判。对于由LLM驱动的机器人的治理基准测试,其中仿真器证明准入/策略/合同/审计流程行为正确,物体交接(抓取、携带、放置)时的接触保真度成为一种负担:接触力积分噪声引入了审计链分歧,该分歧与被测治理属性在结构上无关。我们提出有界保真度模拟即演示阶段,这是一种设计模式,在明确括起的交接包络内抑制接触物理,同时在其他地方保持完整动力学。该构造使用由220行Python适配器驱动的MuJoCo的动捕体原语,治理桥通过结构化意图调用该适配器。我们将审计链稳定性形式化为重放中哈希事件日志的字节相等性,并识别出两个隐含它的结构包络属性。在每种姿态N=1000次重放中,动捕变体产生一个不同的审计链哈希(1000/1000字节相同;Wilson 95% CI [0.997, 1.000]);接触力基线产生584个不同哈希(993/1000分歧;CI [0.987, 0.998])。时间步扫描(1、2、5、10毫秒)表明分歧是结构性的,而非调优伪影:在每个时间步它保持在0.985。包络边缘时间抖动(+/-10仿真步,1,400次重放)产生0分歧,并且审计链在K在{1, 2, 3}中顺序交接的物体(1,500次重放)中保持字节相等,每次拾放开销亚线性。该模式以近乎零的工程成本为基准设计者提供审计链可重现性;我们还绘制了其有害之处(仿真到现实验证、策略训练、接触丰富任务),以免误部署。
英文摘要
Sim-to-real research pursues physics fidelity as a primary objective: simulators are judged by how closely they reproduce real-world contact dynamics. For governance benchmarking of LLM-driven robots, where the simulator demonstrates that an admission/policy/contract/audit pipeline behaves correctly, contact fidelity at object handoffs (grasp, carry, place) becomes a liability: contact-force integration noise injects audit-chain divergence that is structurally unrelated to the governance property under test. We propose bounded-fidelity sim-as-demo-stage, a design pattern that suppresses contact physics within explicitly bracketed handoff envelopes while preserving full dynamics elsewhere. The construction uses MuJoCo's mocap-body primitive driven by a 220-line Python adapter that the governance bridge invokes via structured intents. We formalise audit-chain stability as byte-equality of the hashed event log across replays and identify two structural envelope properties that imply it. Across N=1000 replays per posture, the mocap variant produces one distinct audit-chain hash (1000/1000 byte-identical; Wilson 95% CI [0.997, 1.000]); the contact-force baseline produces 584 distinct hashes (993/1000 diverged; CI [0.987, 0.998]). A timestep sweep (1, 2, 5, 10 ms) shows the divergence is structural, not a tuning artefact: it stays at 0.985 at every timestep. Envelope-edge timing jitter (+/-10 simulation steps, 1,400 replays) produces 0 divergence, and audit chains remain byte-equal across K in {1, 2, 3} sequentially handed-off objects (1,500 replays) with sub-linear per-pick-and-place overhead. The pattern gives benchmark designers audit-chain reproducibility at near-zero engineering cost; we also map where it is harmful (sim-to-real validation, policy training, contact-rich tasks) so it is not mis-deployed.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- Soochow University(苏州大学)
机构由 AI 辅助整理,请以论文原文为准。