arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识迁移参数可以被学习吗?LePoKet 用于高效机器人视觉

Can Knowledge Transfer Parameters Be Learned? LePoKet for Efficient Robotic Vision

Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia

arXiv 2609.16637首次发表:更新:

发表机构

Université Paris-Saclay; IBISC Laboratory; Universidad de La Sabana(巴黎-萨克雷大学; IBISC实验室; 萨瓦纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 LePoKet 框架,通过可学习的结构迁移参数优化,在识别与运动感知任务中显著提升紧凑模型的性能,实现高效机器人视觉。

AI 中文摘要

高效感知是机器人在计算、内存和延迟预算受限条件下运行的核心问题。从更大的预训练模型进行知识迁移,为增强紧凑型感知网络提供了一条实用途径,但现有方法通常依赖于固定的蒸馏目标或手动设计的交互机制。基于遗传知识迁移(HKT),我们提出了 LePoKet(可学习的知识迁移参数优化),这是一种结构迁移框架,将知识继承直接嵌入前向计算中。LePoKet 引入了一个分块的提取-变换-混合接口,其交互参数通过可学习遗传注意力(LGA)算子与子网络联合优化,无需辅助蒸馏损失或温度缩放。我们首先在 CIFAR-10 和 CIFAR-100 上使用 ResNet 父子对对该机制进行了表征,与标准子网络训练相比,相对误差分别降低了 24.57% 和 25.1%。随后,我们将 LePoKet 集成到一个仅在 FlyingChairs 和 FlyingThings3D 上训练的紧凑型 RAFT 光流模型中,用于密集运动估计评估。LePoKet 将紧凑型 RAFT 基线在 Sintel Clean 上的 EPE 从 2.21 提升至 1.92,在 Sintel Final 上从 3.35 提升至 3.01,在 KITTI 上从 7.51 提升至 6.39。与 HKT 的直接比较进一步表明,LePoKet 将 CIFAR-10 准确率从 92.40% 提升至 93.40%,同时在所评估的紧凑型迁移变体中取得了最佳的 Sintel Final 和 KITTI 误差,并在 Sintel Clean 上表现相当。这些结果表明,可学习的结构迁移能够泛化到识别和运动感知任务,并为高效机器人视觉提供了一种有前景的方法。

英文摘要

Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larger pretrained models offers a practical route to stronger compact perception networks, but existing approaches commonly rely on fixed distillation objectives or manually designed interaction mechanisms. Building on Hereditary Knowledge Transfer (HKT), we propose LePoKet (Learnable Parameter Optimization for Knowledge Transfer), a structural transfer framework that embeds knowledge inheritance directly into the forward computation. LePoKet introduces a block-wise Extract-Transform-Mix interface whose interaction parameters are optimized jointly with the child network through a Learnable Genetic Attention (LGA) operator, without auxiliary distillation losses or temperature scaling. We first characterize the mechanism on CIFAR-10 and CIFAR-100 using ResNet parent-child pairs, obtaining relative error reductions of 24.57% and 25.1%, respectively, over standard child training. We then evaluate LePoKet for dense motion estimation by integrating it into a compact RAFT-based optical-flow model trained only on FlyingChairs and FlyingThings3D. LePoKet improves the compact RAFT baseline from 2.21 to 1.92 EPE on Sintel Clean, from 3.35 to 3.01 on Sintel Final, and from 7.51 to 6.39 on KITTI. A direct comparison with HKT further shows that LePoKet improves CIFAR-10 accuracy from 92.40% to 93.40% while achieving the best Sintel Final and KITTI errors among the evaluated compact transfer variants, with comparable performance on Sintel Clean. These results demonstrate that learnable structural transfer generalizes across recognition and motion perception tasks and provides a promising approach for efficient robotic vision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑