发表机构
Iowa State University; BRAC University(爱荷华州立大学; BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对受含义随观测和动作变化约束的工具使用语言模型智能体,提出有状态CARS方法,通过冻结状态-延续模式实现跨历史无效性证书复用,实验表明其精确性高但系统性能未优于官方CARS。
AI 中文摘要
使用工具的语言模型智能体面临含义随观测和先前动作变化的约束。我们研究在有状态验证器(hard stateful validator)约束下,利用跨历史的无效性证书(invalidity certificates)从模型分布进行精确采样。有状态CARS(Stateful CARS)在每次尝试中冻结一组可靠的状态-延续模式(state-continuation schemas),并移除包含在匹配抽象状态处已验证延续的所有轨迹。通过精确的剩余杜布变换(exact residual Doob transform)从得到的提议分布中采样。我们给出可验证的未来有效性互模拟条件,证明模式的可靠性、自适应精确性、独立同分布输出、几乎必然终止、单调接受性和压缩不变性,并通过可达完整历史乘积状态的数量来表征计算量。对于依赖历史的语言模型,该数量可能呈指数级增长;因此,所评估的方法不做通用有限前缀树(finite-trie)可扩展性的断言。在可枚举工作流中,其解析律在有效性概率为6×10⁻⁸时与有效条件分布的匹配度达10⁻¹⁶,而感知状态的局部解码可能偏离0.97。一项匹配比较结果为负面:观测键控的官方CARS在采样器步数上更廉价(根/有状态的比率为0.942 [0.934,0.951]),与通义千问(Qwen)的比较无差异(0.99 [0.90,1.08]);仅在内部匹配键消融实验中,跨历史转移有帮助(1.27倍)。因此,证据支持基于模式的精确条件化,而非相对于CARS的系统优势。
英文摘要
Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certificates across histories. Stateful CARS freezes a bank of sound state--continuation schemas within each attempt and removes every trajectory containing a certified continuation at a matching abstract state. An exact residual Doob transform samples from the resulting proposal. We give a checkable future-validity bisimulation condition, prove schema soundness, adaptive exactness, i.i.d.\ outputs, almost-sure termination, monotone acceptance, and compression invariance, and characterize computation by the number of reachable full-history product states. This number can be exponential for a history-dependent language model; the evaluated method therefore makes no generic finite-trie scalability claim. On enumerable workflows, its analytic law matches the valid conditional to $10^{-16}$ at validity probability $6\times10^{-8}$, whereas state-aware local decoding can be $0.97$ away. A matched comparison is negative: observation-keyed official CARS is cheaper in sampler steps (root/Stateful ratio $0.942$ $[0.934,0.951]$), and the Qwen comparison is null ($0.99$ $[0.90,1.08]$). Cross-history transfer helps only in an internal matched-key ablation ($1.27\times$). Thus the evidence supports exact schema-induced conditioning, not a systems advantage over CARS.