arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Paritok-4B:面向编码智能体的意图条件上下文压缩模型

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Jiayu Shi, Luzhuo Chen

arXiv 2608.24188首次发表:更新:

发表机构

Paritok(Paritok)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对编码智能体上下文占比过高的问题,提出意图条件抽取式压缩模型 Paritok-4B,在 SWE-bench Lite 上实现高压缩率的同时,未显著降低求解率,且成本更低、可自托管。

AI 中文摘要

编码智能体每一轮都会将大型文件读取结果和工具输出重新发送给前沿大语言模型,这类上下文占据了其 token 账单的主要部分。通用提示压缩器是基于散文文本训练的,不太适合代码场景:它们会改写标识符并丢失智能体编辑所需的精确文本片段。我们提出 Paritok-4B,这是一款基于两项核心原则构建的、面向编码智能体轨迹的 4B 规模 LoRA 压缩器。它是抽取式的:会直接选择文本片段而非重写,其输出的标识符、路径和数字中有 96.0% 已出现在输入中,在保留的 SWE-bench Lite 输出上这一比例为 96.2%。它是意图条件的:在获知智能体当前任务后,主要在保留的片段内操作,选择哪些行保留(保留的行比移除的行与意图的相关性高 0.067,配对 95% 置信区间为 [+0.056, +0.078]),而非改变保留内容的多少。我们在 67074 条真实 OpenHands 轨迹上蒸馏 gpt-4.1-mini 教师模型,得到 40606 个经验证的示例,对 Qwen3-4B 进行微调。在全部 300 个 SWE-bench Lite 实例上,Paritok-4B 将智能体上下文压缩至原大小的 25.7%,压缩难度是 gpt-4.1-mini 压缩器(压缩至原大小的 50.2%)的 2.0 倍,是 gpt-5(压缩至原大小的 61.9%)的 2.4 倍,同时保留了未压缩单次求解质量的 86.5%。输入真实智能体生成的带 -n 行号的 cat 命令输出时,它的压缩率略低(27.8%),但保留的求解质量更高(89.3%);此时配对检验具有统计学意义,有 30 个实例仅通过未压缩上下文求解,17 个仅通过压缩上下文求解,精确 McNemar 检验 p 值为 0.079,因此在该样本规模下,将上下文压缩至原大小的约四分之一并不会显著降低求解率。该模型是一个 264 MB 的适配器,可在单块 24 GB GPU 上自托管,无需按 token 收取压缩器费用,按标价计算这决定了其经济性:使用 gpt-5 作为压缩器会产生净成本,其花费超过了它所节省的下游 token 成本。模型权重、数据和评估脚本均以 Apache 2.0 许可开源。

英文摘要

Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code poorly: they paraphrase identifiers and drop the exact spans an agent needs to edit. We present Paritok-4B, a 4B LoRA compressor for coding-agent trajectories built on two commitments. It is extractive: it selects spans rather than rewriting them, and 96.0% of the identifiers, paths, and numbers it emits already appear in its input, holding at 96.2% on held-out SWE-bench Lite output. It is intent-conditioned: told the agent's current task, it acts chiefly inside a retained segment, selecting which lines survive (retained lines are +0.067 more intent-relevant than removed ones, paired 95% CI [+0.056, +0.078]) rather than changing how much is retained. We distil a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories into 40,606 validated examples and fine-tune Qwen3-4B. On all 300 SWE-bench Lite instances, Paritok-4B compresses agent context to 25.7% of its size, 2.0x harder than a gpt-4.1-mini compressor (50.2%) and 2.4x harder than gpt-5 (61.9%), while retaining 86.5% of uncompressed single-shot solve quality. Fed the cat -n line-numbered input real agents produce, it compresses slightly less (27.8%) and retains more (89.3%); there the paired test is informative, with 30 instances solved only uncompressed and 17 only compressed, an exact McNemar p=0.079, so at this sample size compressing context to roughly a quarter of its size does not significantly reduce the solve rate. The model is a 264 MB adapter that self-hosts on one 24 GB GPU with no per-token compressor fee, which at list prices decides the economics: gpt-5 as a compressor is net-negative, costing more than the downstream tokens it saves. Weights, data, and evaluation scripts are open (Apache 2.0).

Comments20 pages, 1 figure, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑