Textual Entailment is not a Better Bias Metric than Token Probability
文本蕴含不是比标记概率更好的偏见度量标准
机构 * Information Sciences Institute University of Southern California(信息科学研究所 美国南加州大学)
AI总结 本文通过实验表明,自然语言推理(NLI)与标记概率(TP)在评估语言模型偏见时表现不同,NLI更不稳定且对刻板印象句子更敏感,因此不推荐作为TP的替代方案。
Comments 12 pages, 1 figure. Substantial revisions following October 2025 ARR Cycle. Currently under review in January 2026 ARR Cycle