arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32547cs.LG

优先进行重复LLM评估以发现隐藏故障

Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery

Keita Broadwater, Akin Broadwater

首次发表
浏览论文内容

中文总结 AI 辅助

针对浅层评估遗漏低概率故障的问题,提出预算化发现框架,利用浅层结果与提示表示排序零故障提示以优先深度评估,在AIRBench上实现最高2.54倍隐藏故障发现提升。

中文摘要 AI 辅助

大型语言模型通常通过在基准测试中为每个提示生成少量随机响应来进行评估。由于推理预算有限,这种浅层评估可能无法观察到低概率但操作上重要的故障。因此,在少量样本中未产生故障的提示可能看似可靠,尽管在重复推理下其潜在故障概率非零。我们将LLM可靠性评估构建为一个预算受限的发现问题,其中每个提示与一个未知的每次生成故障概率相关联。我们提出了一种预算化发现框架,该框架首先对提示集进行浅层评估,然后利用试验级故障结果以及提示派生的表示来学习基于特征的故障倾向排名。由此产生的分数优先对零观测浅层故障的提示进行深度评估,将深度评估预算集中在更可能发现隐藏故障的地方。我们在AIRBench和StrongREJECT上,在多种模型和系统提示条件下评估了这种方法。核心实证测试询问:未访问深度评估结果而训练的模型,能否根据零观测浅层故障的提示在深度评估下产生故障的可能性对其进行排名。在AIRBench上,排名前10%的未解决提示对Qwen 2.5 7B实现了2.54倍的隐藏故障提升,对Gemma 3n E4B实现了1.87倍的提升,分别恢复了随后观察到的隐藏故障的25.4%和18.7%,而随机分配下的预期为10%。语义邻域和特征消融分析进一步表明,这种预测信号可以从提示内容和关系的多种表示中恢复。

英文摘要

Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failures in a small sample may therefore appear reliable despite having a nonzero latent probability of failure under repeated inference. We formulate LLM reliability evaluation as a budget-constrained discovery problem in which each prompt is associated with an unknown per-generation failure probability. We propose a budgeted discovery framework that first performs shallow evaluation across the prompt set and then uses trial-level failure outcomes together with prompt-derived representations to learn a feature-based ranking of failure propensity. The resulting scores prioritize prompts with zero observed shallow failures for deeper evaluation, concentrating the deep-evaluation budget where hidden failures are more likely to be discovered. We evaluate this approach on AIRBench and StrongREJECT across multiple model and system-prompt conditions. The central empirical test asks whether models fit without access to deep-evaluation outcomes can rank prompts with zero observed shallow failures according to their likelihood of producing failures under deeper evaluation. On AIRBench, the highest-ranked 10\% of unresolved prompts achieves 2.54x hidden-failure lift for Qwen 2.5 7B and 1.87x for Gemma 3n E4B, recovering 25.4\% and 18.7\% of subsequently observed hidden failures, respectively, compared with 10\% expected under random allocation. Semantic-neighborhood and feature-ablation analyses further show that this predictive signal can be recovered from multiple representations of prompt content and relationships.

补充信息

↑