arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13235cs.ROcs.CV

通过结果瓶颈学习操作充分的表示

Learning Manipulation-Sufficient Representations via Outcome Bottlenecks

  • Yeungnam University(岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

Md Selim Sarowar, Sungho Kim

AI总结:

针对网络化操作中带宽受限问题,提出基于结果瓶颈的随机表示学习方法,以最小统计量保留动作结果分布,在抓取等任务上显著优于重建几何表示,并大幅降低通信开销。

AI中文摘要:

网络化的操作端点将感知与驱动连接在计算和带宽受限的链路上,然而它们通常交换为保真度而非动作结果优化的密集几何状态。我们通过一种无策略、动作条件的结果瓶颈学习随机表示:边际结果对数损失提供失真项,KL项对码率进行正则化。该构造的动机是保留每个允许动作的结果分布的最小统计量,而实现的有限模型被评估为码率正则化的混合预测器。相同的编码器和结果头支持抓取选择、单例共形滤波、主动视点选择和潜在测试时自适应。一个有限探针定理将局部水平集的切空间与结果雅可比矩阵的零空间等同起来。合成预言机验证了这一结果;在扫描对象上,解析替代与模拟器测量的不变性在0.56°内一致。在13个对象上的11,979次模拟抓取中,重建几何的力旋量分数对提升成功率的AUC为0.542,在弯曲对象上低于随机水平,而我们的表示达到0.876。在25%的承诺水平下,执行抓取成功率为0.503对0.984。512字节的接口比一帧RGB-D图像小288倍,每次CPU决策运行时间为16毫秒。独立的合成点跟踪测试的共形水平;扫描对象的全对覆盖率作为聚类经验诊断报告。在未见对象上,场景内AUC降至0.569,而全反馈更新将经验平均成对覆盖率从0.728提高到0.883。

英文摘要:

Networked manipulation endpoints couple perception to actuation across compute- and bandwidth-limited links, yet commonly exchange dense geometric states optimized for fidelity rather than action outcomes. A stochastic representation is learned with a policy-free, action-conditioned outcome bottleneck: marginal outcome log-loss supplies distortion and a KL term regularizes rate. The construction is motivated by the minimal statistic that preserves the outcome distribution of every admissible action, while the implemented finite model is evaluated as a rate-regularized mixture predictor. The same encoder and outcome head support grasp selection, singleton conformal filtering, active viewpoint selection, and latent test-time adaptation. A finite-probe theorem identifies the local level-set tangent space with the null space of an outcome Jacobian. The synthetic oracle verifies this result; on scanned objects, an analytic surrogate agrees with measured simulator invariances within \(0.56^\circ\). Across 11,979 simulated grasps on 13 objects, a reconstructed-geometry wrench score attains 0.542 AUC against lift success and falls below chance on curved objects, while our representation attains 0.876. At 25\% commitment, executed-grasp success is 0.503 versus 0.984. The 512-byte interface is \(288\times\) smaller than one RGB-D frame and runs at 16\,ms per CPU decision. Independent synthetic points track the tested conformal levels; scanned-object all-pair coverage is reported as a clustered empirical diagnostic. On unseen objects, within-scene AUC falls to 0.569, and a full-feedback update raises empirical mean pairwise coverage from 0.728 to 0.883.

补充信息

↑