arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带符号词汇置信度用于风险校准的意图路由

Signed Lexical Confidence for Risk-Calibrated Intent Routing

Yezhou Cheng, Zehua Yang, Bojun Lin

arXiv 2610.00262首次发表:更新:

发表机构

University of Wisconsin–Madison; Pinterest Inc.(威斯康星大学麦迪逊分校; Pinterest 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种带符号词汇置信度门控,结合分类器对数几率与词汇支持,通过二项校准优化风险阈值,在多个意图数据集上显著降低风险覆盖曲线面积并提升接受覆盖率。

AI 中文摘要

选择性意图路由允许助手在可靠预测上采取行动,同时推迟处理不确定的请求。标准置信度分数主要反映基础模型的表示,这留下了在不改变其决策的情况下纳入补充证据的机会。我们引入了一个带符号的词汇门控,它将句子分类器的对数几率边际与稀疏词汇模型对分类器预测意图的支持相结合。通过将正证据分配给词汇一致性,将负证据分配给词汇上占优势的竞争意图,该门控比无符号词汇置信度或硬性一致性规则保留了更多信息。一个独立的二项校准阶段为指定的风险目标选择操作阈值。在BANKING77、CLINC150和HWU64上的十次运行中,所提出的分数相对于仅学习的语义门控,将风险覆盖曲线下的面积分别减少了15.8%、15.1%和11.8%。在名义5%错误目标下,它在BANKING77和HWU64上分别将接受覆盖率提高了1.83和5.14个百分点,而CLINC150已接近完全覆盖。在更严格的2%目标下,同时二项程序在可用的校准预算下,在所有30个数据集-运行组合中产生了非空策略。匹配对照表明,所提出的特征在平均错误排名上优于测试的无符号词汇置信度特征,并在数据集依赖的情况下优于二元一致性。所得的双特征门控为风险校准的意图路由提供了紧凑、可解释的置信度增强,同时保留了基础分类器的预测。

英文摘要

Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests. Standard confidence scores primarily reflect the base model's representation, leaving an opportunity to incorporate complementary evidence without changing its decisions. We introduce a signed lexical gate that combines a sentence classifier's logit margin with a sparse lexical model's support for the classifier's predicted intent. By assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent, the gate retains more information than either unsigned lexical confidence or a hard agreement rule. An independent binomial calibration stage selects an operating threshold for a specified risk target. Across ten runs on BANKING77, CLINC150, and HWU64, the proposed score reduces area under the risk-coverage curve by 15.8%, 15.1%, and 11.8% relative to a learned semantic-only gate. At a nominal 5% error target, it increases accepted coverage by 1.83 and 5.14 percentage points on BANKING77 and HWU64, while CLINC150 is already near full coverage. At a stricter 2% target, the simultaneous binomial procedure yields a nonempty policy in all 30 dataset-run combinations at the available calibration budgets. Matched controls show that the proposed feature improves average error ranking over the tested unsigned lexical-confidence feature, with dataset-dependent gains over binary agreement. The resulting two-feature gate provides a compact, interpretable confidence enhancement for risk-calibrated intent routing while preserving the base classifier's predictions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑