arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03387cs.AIcs.CL

类型化决策模型中候选覆盖率的基准测试

Benchmarking candidate coverage and rejection policy transfer in typed decision models

Jiawen Lu, Tongtong Wu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出配对候选覆盖率基准协议,评估Laya和Jev在四个数据集上的缺失答案检测与有效候选拒绝能力,发现原生拒绝行为差异显著,并强调分类、排名和拒绝策略需分别测量。

中文摘要 AI 辅助

类型化决策模型在请求时返回对所提供的答案选项的选择或分布。在完整选项下的准确性并不能确定模型是否识别出参考答案缺失,或避免拒绝有效候选。我们提出了一种配对的候选覆盖率基准协议,并对Laya和Jev在AG News、DBpedia、Emotion和TREC数据集上进行了初步评估。模型接收相同的冻结文本和请求:300个校准文本和589个测试文本,每个模型产生23,932个预测。存在/缺失对匹配普通候选数量,名称变体保留描述、成员和顺序。原生拒绝行为差异显著:在TREC的五个自然名称候选下,Laya检测出97.2%的缺失答案案例,但错误拒绝了69.7%的存在控制案例;Jev的比率分别为24.8%和0.0%。仅校准的无分数阈值将这些比率分别改变为33.9%/3.7%和45.0%/1.8%。在DBpedia上,Jev的高覆盖率分数AUROC支持更强的操作点,而两个模型在Emotion上的完整集准确性均较弱。能力条件分析、概率精度敏感性和接口审计表明,分类、分数排名和拒绝策略需要分别测量。此初步基准是描述性的,仅限于参考标签遗漏;它未确立自然范围外泛化、因果机制或新的拒绝方法。

英文摘要

Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference labels, testing Laya/Jev policy transfer at equal calibration budgets and all three models on artificial omission, natural retrieval misses and public out-of-scope queries. Source calibration often fails to preserve the target operating point. A Jev policy calibrated on DBpedia rejects 69.3% of covered Emotion test inputs, while an Emotion policy loses detection entirely. Retrieval exposes a different tradeoff: with ten intent candidates, Laya detects 99.0% of out-of-scope queries but rejects 48.8% of covered queries. Separating missing-answer sources reveals these costs alongside retrieval coverage. The benchmark provides shared inputs, explicit decision and failure categories, and reproducible scoring to assess rejection policies under the conditions in which they are reused. Code and benchmark artifacts are available at https://github.com/luckykevvv/Decision_Model_Benchmark.

发表机构

  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑