arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习的纳米四旋翼动力学中的评估记录污染:新种子审计

Evaluation-Recording Contamination in Learned Nano-Quadrotor Dynamics: A Fresh-Seed Audit

David Shulman

arXiv 2607.18482首次发表:更新:

AI 中文总结

研究学习的纳米四旋翼动力学中评估记录污染问题,通过固定NanoBench Crazyflie 2.1数据集快照进行实验,比较不同模型,发现污染虽降误差但置信区间过零,可靠评估需记录级分割等多方面操作。

AI 中文摘要

学习的飞行动力学模型通常在从较长记录中提取的短窗口上进行训练。随机分割这些窗口可能会将来自同一物理飞行的相关样本同时放入训练集和评估集。我们使用NanoBench Crazyflie 2.1数据集的固定快照来审计此问题。实验中评估飞行、展开起始、训练窗口数量、验证数据和优化预算保持固定。在污染协议中,[LeakagePercent]百分比的不相交训练集被评估记录中的窗口替换,双臂使用相同的干净验证集。我们比较了世界坐标增量多层感知器和相对坐标控制,跨越[NumSeeds]个新训练种子和[NumEvalFlights]个完整评估记录。主要终点是1秒展开上的故障感知位置均方根误差,通过配对交叉记录与种子推理进行分析。污染使表观世界模型误差降低了12.1%,从0.299米降至0.262米,但污染减去不相交误差的预先声明的95%置信区间为-0.0735至0.0011米。由于区间穿过零,验证结果为阴性。次要视野和相对坐标模型显示出类似趋势,但不支持坐标表示交互。因此,可靠的泄漏评估需要记录级分割、故障感知展开、跨新种子复制以及对记录和优化种子进行推理。

英文摘要

Learned flight-dynamics models are often trained on short windows extracted from longer recordings. Randomly splitting these windows can place dependent samples from the same physical flight in both training and evaluation sets. We audit this issue using a fixed snapshot of the NanoBench Crazyflie 2.1 dataset. The experiment holds evaluation flights, rollout starts, training-window count, validation data, and optimization budget fixed. In the contaminated protocol, [LeakagePercent] percent of a recording-disjoint training set is replaced with windows from evaluation recordings, while both arms use the same clean validation set. We compare a world-coordinate delta multilayer perceptron with a relative-coordinate control across [NumSeeds] fresh training seeds and [NumEvalFlights] complete evaluation recordings. The primary endpoint is failure-aware position RMSE over 1-second rollouts, analyzed with paired crossed recording-by-seed inference. Contamination lowers apparent world-model error by 12.1 percent, from 0.299 m to 0.262 m, but the predeclared 95 percent confidence interval for contaminated minus disjoint error is -0.0735 to 0.0011 m. Because the interval crosses zero, the confirmatory result is negative. Secondary horizons and the relative-coordinate model show similar trends, but no coordinate-representation interaction is supported. Reliable leakage assessment therefore requires recording-level splits, failure-aware rollouts, replication across fresh seeds, and inference over both recordings and optimization seeds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑