增值税判定(VAT Determination)中大型语言模型智能体(LLM-Agent)分解的规模适配:一项试点控制扫描研究
Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep
浏览论文内容
中文总结 AI 辅助
本试点研究在带反向收费的跨境增值税判定场景中,对比不同LLM智能体分解配置,发现中间配置准确率领先但未达标准,提出分解规模适配启发式方法并发布相关资源。
中文摘要 AI 辅助
近期的大型语言模型智能体(LLM-Agent)系统存在相互冲突的设计选择:是将工作分解给多个窄域智能体,还是使用一个强大的工具型智能体。本试点研究针对带反向收费的受限跨境增值税(VAT)判定场景开展该选择的研究,该场景中每个案例都有一个 oracle 标签,且每个中间决策均可独立评分。我们固定活动表面(包括子任务、工具、输入输出模式、验证检查、编排器、基础模型及合并策略),仅改变子任务在四个编排配置中分配给工作者的方式,从1个宽域工作者到5个窄域工作者,对比S0(一个经调优的无编排器单智能体),采用确定性规则引擎作为oracle。该程序共包含4400次运行:40个案例、5次重复的主扫描,匹配标记分支用于分离提示预算与智能体数量的影响,以及3个故障注入分支,所有运行均按预先注册的证伪标准评判。两个中间配置的准确率领先(分别为0.830,而端点配置为0.720和0.770),但未达到与精细端点对比的预先设定标准,因此在试点规模下,中间最优假设未得到支持。单智能体未在帕累托意义上优于编排集。匹配标记标准生效:预算匹配的单智能体比领先者低6.5个百分点,但区间包含零,因此任何优势均与提示预算的解释一致。在故障注入场景下,可用性故障在所有粒度级别均被吸收,宽范围重启将其基线过度恢复了+0.160,而一条符合模式的幻觉记录会降低所有配置的性能并反转排序,对碎片化配置的影响最严重。本研究的贡献是一种受限的、预先注册的分解规模适配试点启发式方法(在依赖层中点设置一个分区边界),并附带oracle、数据集、工具链、原始轨迹和分析流程发布。
英文摘要
Recent LLM-agent systems make conflicting design bets: decompose work across many narrow agents, or use one strong tool-using agent. This pilot studies that choice on bounded cross-border VAT determination with reverse charge, where every case has an oracle label and each intermediate decision is independently scoreable. We hold the activity surface fixed (subtasks, tools, I/O schemas, validation checks, orchestrator, base model, and merge policy) and vary only the assignment of subtasks to workers across four orchestrated configurations, from one wide worker to five narrow ones, against S0, a tuned no-orchestrator single agent, with a deterministic rule engine as oracle. The program spans 4,400 runs: a 40-case, five-repeat main sweep, matched-token arms separating prompt-budget from agent-count effects, and three failure-injection arms, all judged against pre-registered falsification criteria. The two intermediate configurations lead on accuracy (0.830, against endpoints at 0.720 and 0.770) but miss the pre-stated bar against the fine endpoint, so the intermediate-optimum hypothesis remains unsupported at pilot scale. The single agent does not Pareto-dominate the orchestrated set. The matched-token criterion fires: the budget-matched single agent lands 6.5 points below the leader, but the interval includes zero, so any advantage is consistent with a prompt-budget explanation. Under injection, availability faults are absorbed at every granularity, with wide-scope restart over-recovering its baseline by +0.160, while one schema-conforming hallucinated record degrades every configuration and inverts the ordering, hitting fragmented configurations hardest. The contribution is a bounded, preregistered pilot heuristic for right-sizing decomposition (place one partition boundary at the dependency-layer midpoint), released with oracle, dataset, harness, raw traces, and analysis pipeline.