arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04629cs.AIcs.LGcs.SYeess.SY

SiLR:面向大语言模型工具智能体的结构保留准入与过程奖励

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

  • School of Engineering, Institute of Science Tokyo(科学东京学院工程学院)
  • College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院)
  • Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou

中文总结 AI 辅助

该研究针对LLM工具智能体提出SiLR方法,通过乘积顺序的结构保留准入解决标量投影陷阱,在多场景基准中恢复表现显著,作为GRPO奖励也优于基线,证明结构保留的乘积顺序是解决违规问题的充分方案。

中文摘要 AI 辅助

大语言模型(LLM)工具智能体的运行时门通常被视为一种过滤器。在ReAct循环中,被拒绝的提议将在相同状态下被另一个提议取代,因此该门是对提议流的搜索算子,其准入准则决定了哪些轨迹是可达的。我们研究违规后恢复准入,即系统仍处于违规状态时必须允许进展,并识别出标量投影陷阱:聚合得分门会接受局部改进的提议,使轨迹陷入平台期。SiLR则对每个提议进行影子执行,并基于分支级违规状态(过载分支支持和每分支严重程度)的乘积顺序准入。我们证明,对于该顺序,没有标量替代方案是可靠的,因此这种失败是表征层面的问题,而非阈值调整的问题。在挖掘的Gym-ANM场景中,SiLR在21/21多动作回合中实现了恢复,而终端门和最佳标量门分别仅为0/21和9/21,在全部24个场景基准中具有显著优势。终端与结构化的二分性在三个模型家族及CityLearn中均成立。由于准入依赖确定性模拟,LLM处于信任边界之外:一种能同时击败标量门和仅支持基线的幅度重分布攻击,仅被完整的每分支谓词所遏制。当两个约束家族活跃时,所有测试的标量投影都会准入物理不安全的动作;仅支持基线的准入比例最大(42410个中的63.2%;乘积顺序为0)。在最困难的双家族轨迹中,标量门仅通过该不安全类别实现恢复。作为GRPO过程奖励复用后,其在所有挖掘场景中均优于计数投影,且是唯一测试的奖励,其无门策略超过未训练的基线(0.844 vs. 0.778)。标量投影在两个设计点均丢失了违规几何结构;只有完整的乘积顺序在结构上是充分的。

英文摘要

A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shapes which trajectories are reachable. We study post-violation recovery admission, where progress must be admitted while the system is still in violation, and identify the scalar projection trap: an aggregate-score gate accepts a locally improving proposal and commits the trajectory to a plateau. SiLR instead shadow-executes each proposal and admits it under a product order over the branch-level violation state (overloaded-branch support and per-branch severity). We prove that no scalar surrogate is sound for this order, so the failure is representational, not a matter of threshold tuning. On mined Gym-ANM scenarios, SiLR recovers 21/21 multi-action episodes against 0/21 for terminal and 9/21 for the best scalar gate, significant across the full 24-scenario benchmark. The terminal-versus-structured dichotomy holds across three model families and in CityLearn. Because admission rests on deterministic simulation, the LLM lies outside the trust boundary: a magnitude-redistribution attack that defeats both scalar and support-only baselines is contained only by the full per-branch predicate. With two constraint families active, every tested scalar projection admits physically unsafe actions; support-only admits the largest fraction (63.2% of 42,410; product order 0). In the hardest dual-family traces, scalar gates recover only through that unsafe class. Reused as a GRPO process reward, it outperforms its count projection in every mined scenario and is the only tested reward whose ungated policy exceeds the untrained base (0.844 vs. 0.778). Scalar projection loses the violation geometry at both design points; only the full product order is structurally sufficient.

补充信息

↑