AI 中文总结
本文通过五组消融实验等方法,对比 LSP 与 grep 检索的 token 效率,发现 LSP 通常不节省 token,仅对最弱模型有 token 节省效果,需根据任务、模型等选择检索工具。
AI 中文摘要
编码智能体的上下文预算大部分花费在检索上。词汇检索(grep)是通用、即时且零配置的,但存在噪声:它无法区分定义、调用和注释。通过语言服务器协议(LSP)进行的语义检索是精确且带类型的,但需要一个正在运行的、已索引的服务器,并为每个符号支付一次往返开销。我们发现,“语义检索更节省 token”这一说法几乎随处可见,但几乎没有被测量过:没有公开来源在智能体任务成功率相同的情况下,分离出 LSP 与词汇检索的 token 差异。本文用一个指标(成功所需 token 数)将该问题形式化,指定了一个五组消融实验以分离语义检索与混淆因素,将三种预先陈述的失败模式映射为可测量变量,并报告了一项初步研究(涉及 Python 和 TypeScript 代码库;模型为 Claude Opus 4.8、Sonnet 4.6、Haiku 4.5)。答案是有条件的,通常为否定。在符号命名的定位任务中,LSP 会增加 token 开销(+6% 至 +118%),且智能体在免费时会忽略它;在引用完整性任务中,它能提供精确性,但无法节省 token,也无法提高智能体彻底性设定的召回上限,仅对最弱的模型能节省 token。工具选择取决于任务:智能体默认在定位任务中使用 grep(语义使用占比 0-6%),但在引用任务中会主动使用 LSP,占比约一半。在通过实际测试执行评分的编辑任务中,差距最为明显:grep 能完美解决多文件重命名问题,仅进行定位的 LSP 会因遗漏调用站点而在四分之三的任务中失败,即使是完整、索引预热、文本丰富的 LSP(每个引用的行内文本,如生产级 LSP-MCP 服务器的做法)也能弥补大部分差距,但无法完全消除,因为重命名必须触及语义引用未包含的注释和字符串。这一结论并非主张 LSP 始终适用,而是应根据任务类别、模型能力和词汇噪声配置自适应路由。
英文摘要
Coding agents spend most of their context budget on retrieval. Lexical retrieval (grep) is universal, instant, and zero-setup, but noisy: it cannot tell a definition from a call from a comment. Semantic retrieval via the Language Server Protocol (LSP) is precise and typed, but needs a running, indexed server and pays a per-symbol round-trip. The claim that semantic retrieval is more token-efficient is, we find, asserted almost everywhere and measured almost nowhere: no public source isolates the LSP-vs-lexical token delta for an agent at equal task-success. This paper formalizes the question with one metric (tokens-to-success), specifies a five-arm ablation isolating semantic retrieval from confounds, maps three pre-stated failure modes onto measurable variables, and reports a preliminary study (Python and TypeScript repos; Claude Opus 4.8, Sonnet 4.6, Haiku 4.5). The answer is conditional and usually negative. On symbol-named localization the LSP costs tokens (+6% to +118%) and the agent ignores it when free. On reference-completeness it buys precision but not token savings and cannot raise the recall ceiling set by agent thoroughness; it saves tokens only for the weakest model. Tool choice is task-dependent: models default to grep on localization (0-6% semantic use) but reach for the LSP about half the time on reference tasks, unprompted. On edits scored by real test execution the gap is starkest: grep solves multi-file renames perfectly, a location-only LSP fails three-quarters of them by missing a call site, and even a complete, index-warmed, text-enriched LSP (each reference's line inline, as production LSP-MCP servers do) recovers most of the gap but cannot close it, since a rename must touch comments and strings that semantic references exclude. The implication is not LSP-always but an adaptive router keyed on task class, model capability, and lexical noise.
Comments13 pages, 6 figures. Code and data: https://github.com/Poytr1/lsp-vs-grep-token-study