以沉默取胜:大语言模型计划评估中的删除非单调性、自主利用和类型状态门控
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation
AI总结:
研究大语言模型计划评估中删除非单调性等问题,通过命题1给出得分变化公式,经实验发现优化器能找到改进结构,GATE可约束搜索,PCSC能消除事后省略拼接,但GATE不验证策略质量。
AI中文摘要:
计划评估者可能会奖励一个变得不那么明确的战略计划。本文研究了大语言模型生成的风险投资路线的分段期望值评分器中的失败情况。命题1给出了在重新定位其前驱并保留下游值时删除内部转换的得分变化:Δ_k = (∏_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]。在一个固定的26条路线群组上,所有57个可接受的删除都与分析恒等式和阈值符号匹配,并且每条路线至少有一个提高得分的删除。一个寻求得分的优化器,被允许重组路线但未被告知利用机制,在26条路线中的21条中发现了超越基线的未发现结构。GATE对26条沉默路线拒绝发布分数,其中26条无诚实暂停情况;拒绝后,54个后续修订中的47个修复为覆盖结构,严格的覆盖改进从26条中的1条增加到13条。一个自适应的编译器感知合著者暴露了注册表来源边界:在所有四个v1/v1.5条件下,义务通道逃避保持在6/6,而增量索引成本下限将击败诚实路线从6/6减少到3/6,通过沉默的可融资性从5/6减少到0/6,但未建立语义完整性。如果一个计划仅因为省略了必要工作而得分更高,那么该计划并未改进;评估产生了省略激励。PCSC检测并消除模型介导的类型状态记录上的事后省略拼接。在测试的合作环境中,GATE充当确定性的搜索塑造约束,而不仅仅是事后过滤器。它不验证任意大语言模型生成策略的语义完整性或现实世界质量。
英文摘要:
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57 admissible deletions matched the analytic identity and threshold sign, and every route had at least one score-improving deletion. A score-seeking optimizer, allowed to restructure routes but not told the exploit mechanism, found baseline-beating uncovered structures in 21/26 routes. GATE refused score release for 26/26 silenced routes with 0/26 honest suspensions; after refusal, 47/54 next revisions repaired to a covered structure, and strict covered improvement rose from 1/26 to 13/26. An adaptive compiler-aware co-author exposed the registry-provenance boundary: obligation-channel evasions remained 6/6 across all four v1/v1.5 conditions, while delta-indexed cost floors reduced beat-honest routes from 6/6 to 3/6 and fundability-by-silence from 5/6 to 0/6 without establishing semantic completeness. If a plan scores better only because it omits necessary work, the plan did not improve; the evaluation created an omission incentive. PCSC detects and neutralizes post-hoc omission splices over model-mediated typed-state records. In the cooperative setting tested, GATE acts as a deterministic search-shaping constraint, not merely a post-hoc filter. It does not verify the semantic completeness or real-world quality of arbitrary LLM-generated strategies.