arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13545cs.ROcs.LG

运行时增量Transformer用于基于强化学习的自适应控制

Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control

Giansalvo Cirrincione, Adriano Fagiolini

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种运行时增量Transformer机制,在强化学习过程中根据上下文分布有效秩和头输出幅度动态增减注意力头数,消除了长记忆下固定容量控制器的灾难性故障,并在Stribeck摩擦两连杆机械臂上实现全记忆范围成功。

中文摘要 AI 辅助

针对具有不可观测摩擦记忆的机器人机械臂的学习型自适应控制问题,已有研究采用基于注意力的元控制器来解决,此类控制器的注意力头数量在训练前固定,并通过昂贵的离线搜索进行调优。在长记忆时间范围内,这种固定容量的控制器在相当一部分训练种子上容易发生灾难性故障。本文提出了一种运行时机制,在强化学习过程中动态增加和剪枝注意力模块的头数,该机制由两个信号驱动:在线策略上下文分布的有效秩(当表示能力不足时触发增长)以及每个头的输出幅度(标记冗余头以供移除)。本文从理论上分析了增长事件时的策略连续性和剪枝事件时的定量界限。在具有Stribeck摩擦的两连杆机械臂上,所提出的机制在所有记忆范围内均实现了完全成功,消除了长时域故障模式,并免除了对头数进行离线调优的需求。

英文摘要

Learning-based adaptive control of robotic manipulators with non-observable friction memory has been addressed by attention- based meta-controllers whose number of attention heads is fixed before training and is tuned by costly offline search. At long memory horizons, such fixed-capacity controllers are prone to catastrophic failures on a sizeable fraction of training seeds. The present paper introduces a runtime mechanism that grows and prunes the heads of the attention block during reinforcement learning, governed by two signals: the effective rank of the on-policy context distribution, which triggers growth when representational capacity becomes insufficient, and the per-head output magnitude, which flags redundant heads for removal. Policy continuity at growth events and a quantitative bound at prune events are established analytically. On a two- link manipulator with Stribeck friction, the proposed mechanism attains full success across all memory regimes, eliminating the long-horizon failure mode and removing the need for offline tuning of the head count.

发表机构

  • Université de Picardie Jules Verne(皮卡第儒勒·凡尔纳大学)
  • University of Palermo(巴勒莫大学)

机构由 AI 辅助整理,请以论文原文为准。

↑