arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21227cs.RO

FORGE-plus:使用冻结的语言模型监督器进行力预算的富接触装配恢复

FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

Kyupaeck Jeff Rah, Midum Oh

AI总结:

研究在规定力上限下实现紧间隙装配及从插入失败中恢复的问题,提出含冻结大语言模型的两层框架,该模型分配力上限并选恢复策略,通过实验评估,此框架表现良好,解决了部分夹具故障问题,也有负面结果。

AI中文摘要:

力条件强化学习能在规定力上限下实现紧间隙装配,但实际部署需为每个物体确定合适力限制并从不超力的插入失败中恢复。本文提出两层框架,冻结的纯文本大语言模型在执行前为每个物体分配力上限,并使用紧凑文本力签名从固定动作菜单中选择恢复策略。大语言模型不直接控制力,低级控制器执行力上限,恢复策略不能增加它,隐藏的破断力阈值仅评估器知道。在易碎瓶子放置和0.4毫米直径间隙齿轮插入实验中评估该框架,单一策略在易碎和坚固物体上通过256/256评估情节且无破损,正确预测释放时间,完成全表拾取和插入流程,平均峰值力为5.4牛。在注入握内滑动情况下,力签名恢复策略解决了两种夹具40%和64%的故障,而用力更大的基线无效或导致频繁破损。还报告了负面结果,包括近端策略优化算法在严格力约束下无法解决任务以及学习的释放策略未成功。所有实验在刚体模拟中进行,未进行模拟到现实的声明。

英文摘要:

Force-conditioned reinforcement learning (RL) enables tight-clearance assembly under a commanded force ceiling, but practical deployment requires determining an appropriate force limit for each object and recovering from insertion failures without exceeding it. We present a two-layer framework in which a frozen, text-only large language model (LLM) assigns a per-object force ceiling before execution and selects recovery maneuvers from a fixed action menu using compact textual force signatures. The LLM never controls force directly: a low-level controller enforces the force ceiling, the recovery policy cannot increase it, and the hidden breaking-force threshold is known only to the evaluator. We evaluate the framework on fragile bottle placement and 0.4 mm diametral-clearance gear insertion using two grippers (Robotiq 2F-140 and Franka Panda hand). A single policy passes 256/256 evaluation episodes on both fragile and robust objects without breakage, correctly predicts release timing, and completes a full table-pick-and-insert pipeline with a mean peak force of 5.4 N. Under injected in-grip slip, the force-signature recovery strategy resolves 40% and 64% of failures on the two grippers, whereas a press-harder baseline is either ineffective or causes frequent breakage. We also report negative results, including the failure of PPO to solve the task under strict force constraints and unsuccessful learned release strategies. All experiments are conducted in rigid-body simulation with hidden force-threshold breakage; no sim-to-real claim is made.

↑