发表机构
ShanghaiTech University(上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究LLM推理管道中已提交阶段何时值得付出成本,提出受限路径推理(CPR)方法,通过源感知路径假设与阶段级核算结合,在多个实验设置下验证,可测量相关指标,为LLM推理提供新的衡量与验证方式。
AI 中文摘要
在大型语言模型推理管道中,已提交的中间阶段何时值得付出成本?受限路径推理(CPR)将源感知路径假设与阶段级核算相结合。搜索生成临时状态,可信或经过验证的不变量可严格约束,而其他提议则保持灵活且可修订。CPR预测,当任务兼容的承诺的收益超过传播误差和执行成本时,它们可以影响转换、集中候选质量、诱导规律性并暴露反馈。该形式体系涵盖离散承诺和连续流,并测量有效分支、端点集中和每个可用输出的成本。在1180个生成的QCQP和40个工程退化多项式实例(2140个端点)上,残差分类法以17.7%的尝试次数恢复了修复所有问题的额外可行产量的63.0%。固定语言模型核算(跨嵌套分支共享270个唯一调用)发现直接可用产量为41.1%,形式化和确定性执行后为90.0%,单次凸化后为20.0%,全路径为21.1%。在120个配对条件调用中,双动作回滚规则的可用产量达到90%,而反馈条件选择器为36.7%。两个端点探测将源与验证分开:72输出的交叉轨迹移植降低了熵和可接受质量;24输出的同调用自提议试点给出不变的双重复碰撞熵,可用产量分别为25.0%和8.3%,以及1/8的确定性确认端点检查。模型生成的状态提供假设,可信执行获得约束强度。
英文摘要
When does a committed intermediate stage in an LLM reasoning pipeline earn its cost? Constrained Path Reasoning (CPR) pairs a source-aware path hypothesis with stage-level accounting. Search generates provisional states; trusted or validated invariants can constrain hard, while other proposals remain soft and revisable. CPR predicts that task-compatible commitments can factor transitions, concentrate candidate mass, induce regularity, and expose feedback when their gains exceed propagated error and execution cost. The formalism covers discrete commitments and continuous flows and measures effective branching, endpoint concentration, and cost per usable output. Across 1,180 generated QCQPs and 40 engineered degenerate polynomial instances (2,140 endpoints), residual triage recovers 63.0% of repair-all's additional feasible yield with 17.7% of its attempts. Fixed-LLM accounting (270 unique calls shared across nested arms) finds usable yield of 41.1% direct, 90.0% after formalization and deterministic execution, 20.0% after one-shot convexification, and 21.1% for the full path. In 120 paired-condition calls, a two-action rollback rule reaches 90% usable yield versus 36.7% for the feedback-conditioned selector. Two endpoint probes separate source from validation: a 72-output cross-trajectory transplant reduces entropy and acceptable mass; a 24-output same-call self-proposal pilot gives unchanged two-repeat collision entropy, 25.0% versus 8.3% usable yield, and 1/8 deterministically confirmed endpoint checks. Model-generated states supply hypotheses; trusted execution earns constraint strength.