arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双智能体语言模型中继中的状态压缩:约束保留的封闭世界研究

State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation

Anantha Sharma, Sheeba Elizabeth John, Kaarthik Senthil Kumar, Saratsuhas Vijayababu

arXiv 2607.18265首次发表:更新:

AI 中文总结

研究双智能体LLM中继中交接压缩问题,比较无压缩、叙述性总结、JSON提取、基于嵌入的修剪四种交接条件,发现交接表示影响下游可行性,JSON提取可行性准确率最高达0.96,结构化可审计表示利于约束检查。

AI 中文摘要

基于大型语言模型(LLM)的长期运行智能体常常积累包含审计、消除和数值计算的大量中间痕迹。在实践中,这种状态在传递给下游决策步骤前会被压缩,形成信息瓶颈,小的遗漏可能打破严格的数值或分类约束。本文在一个有两个LLM智能体的封闭世界旅行规划中继中评估交接压缩。一名研究者为50个目标实例审计固定的酒店和航班库存,一名预订者仅使用目标和交接有效载荷选择酒店 - 航班对,库存保密。比较了四种交接条件:无压缩、叙述性总结、模式约束的JSON提取和基于嵌入的修剪。对固定库存的穷举枚举提供了精确的可行和最优标签。结果表明,在小决策模型下,交接表示强烈影响下游可行性。JSON提取实现了最高的可行性准确率0.96,而叙述性总结虽产生最小的压缩交接有效载荷,但将可行性降至0.48。基于嵌入的修剪在不进行额外生成压缩调用的情况下,可行性与未压缩控制匹配,为0.88。这些发现表明,约束检查受益于结构化和可审计的交接表示,而非仅依赖简洁性。

英文摘要

Long-running Large Language Model (LLM)-based agents often accumulate large intermediate traces containing audits, eliminations, and numeric calculations. In practice, this state is compressed before handing it to a downstream decision step, creating an information bottleneck in which small omissions can break strict numeric or categorical constraints. This paper evaluates hand-off compression in a closed-world travel-planning relay with two LLM agents. A Researcher audits a fixed inventory of hotels and flights for 50 goal instances, and a Booker selects a hotel--flight pair using only the goal and the hand-off payload, with the inventory withheld. We compare four hand-off conditions: no compression, narrative summarization, schema-constrained JSON extraction, and embedding-based pruning. Exhaustive enumeration over the fixed inventory provides exact feasible and optimal labels. Results show that hand-off representation strongly affects downstream feasibility under a small decision model. JSON extraction achieves the highest feasibility accuracy at 0.96, while narrative summarization, despite producing the smallest compressed hand-off payload, degrades feasibility to 0.48. Embedding-based pruning matches the uncompressed control on feasibility at 0.88 without an additional generative compression call. These findings indicate that constraint checking benefits from structured and auditable hand-off representations rather than relying on brevity alone.

Comments8 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑