arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25018cs.LG

共形级联:多层大语言模型推理的无分布精度保证

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

Yifan Dou, Shikan Lian, Shibo Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对LLM级联推理成本高、置信分数校准不当等问题,提出共形级联框架,以共形预测集大小为推迟规则,提供无分布精度保证,在多基准测试中表现优于启发式级联,且无需模型训练,仅需黑盒API访问。

中文摘要 AI 辅助

大语言模型(LLM)级联通过将简单查询路由到小模型并将困难查询推迟到更大模型来降低推理成本。生产级联通过置信阈值控制这种推迟,但LLM置信分数校准不当,阈值必须针对每个模型对和每个领域进行调整,且没有设置能给出级联精度的形式化界限。我们引入共形级联(CC),这是一个多层推理框架,它使用共形预测集大小作为推迟规则:当校准集收缩到单个答案时接受,否则推迟。该过程提供无分布、有限样本精度保证。通过每层的并集界,接受层的预测集以至少1 - Kα的概率覆盖正确答案;在选择保持条件下,界限收紧到1 - α。我们进一步将预期级联成本表征为α和校准集接受率的显式函数。在跨越科学、医学、常识和标准化考试的18个多项选择基准上,在从四个开放权重模型家族中抽取的两层级联上进行评估,CC在大多数家族 - 基准对上严格优于最强的校准调整启发式级联,在推理繁重的基准上收益最大;在较容易的基准上,级联将绝大多数查询提交给小模型而不损失精度。扩展到开放式生成需要答案聚类步骤,留待未来工作。该方法无需模型训练,仅需黑盒API访问。

英文摘要

Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Production cascades govern this deferral through a confidence threshold, but LLM confidence scores are miscalibrated, the threshold must be tuned per model pair and per domain, and no setting yields a formal bound on cascade accuracy. We introduce \textbf{Conformal Cascade} (CC), a multi-tier inference framework that uses conformal prediction set size as the deferral rule: accept when the calibrated set collapses to a single answer, defer otherwise. The procedure delivers a distribution-free, finite-sample accuracy guarantee. By a per-tier union bound, the prediction set at the accepting tier covers the correct answer with probability at least $1 - Kα$ for any user-specified $α$; under a selection-preservation condition (consistent with, but not strictly implied by, our marginal coverage results), the bound tightens to $1 - α$. We further characterise expected cascade cost as an explicit function of $α$ and the calibration-set acceptance rate. Across 18 multiple-choice benchmarks spanning science, medicine, commonsense, and standardized exams, evaluated on two-tier cascades drawn from four open-weight model families, CC strictly improves over the strongest calibration-tuned heuristic cascade on the majority of family--benchmark pairs, with the largest gains on reasoning-heavy benchmarks where majority vote is unreliable; on easier benchmarks the cascade commits the vast majority of queries to the small model at no accuracy cost. Extension to open-ended generation requires an answer-clustering step that we leave for future work. The method requires no model training and only black-box API access.

发表机构

  • Department of Computer Science, Florida State University(佛罗里达州立大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

↑