评估基于BERT的模型在问答任务上的可靠性
Assessing Reliability of BERT-Based Models on Question Answering Tasks
- Malaviya National Institute of Technology(马拉维亚国家理工学院)
- Central University of Rajasthan(拉贾斯坦中央大学)
- University of Ljubljana(卢布尔雅那大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究评估BERT及其变体在问答任务上的可靠性,发现RoBERTa可靠性更高,ALBERT与DistilBERT存在显著不一致,验证了MCD作为可靠性指标的有效性,强调需同时评估QA模型的准确性与稳定性。
AI中文摘要:
大型语言模型的可靠性估计在很多情况下与其准确性同等重要,因为可靠的模型更值得信赖、更具鲁棒性,且更适合实际应用。自然语言处理(NLP)领域的最新进展,尤其是基于Transformer架构的进展,已显著推动各类NLP任务的发展。本研究聚焦于基于Transformer的问答(QA)模型的可靠性,具体为BERT模型及其变体(RoBERTa、ALBERT、DistilBERT)。这些仅用于编码的预训练Transformer在可视为分类任务的QA任务中展现出卓越的准确性,但其可靠性仍未得到充分探索。本研究通过评估两种条件下的响应稳定性,评估四种基于BERT的模型的可靠性:(1)通过蒙特卡洛Dropout(MCD)引发的内部模型变化;(2)通过意译产生的输入扰动。使用SQuAD和QuAC数据集,我们研究了Dropout率如何影响预测一致性,以及词汇变化是否会影响答案稳定性。研究结果表明,RoBERTa保持更高的可靠性,而ALBERT和DistilBERT则表现出显著的不一致性。统计分析证实,在预测过程中启用MCD不会破坏推理动态,验证了其作为可靠性指标的有效性。这些发现强调,在QA模型中评估准确性和稳定性的重要性,以确保其在实际应用中的稳定性。
英文摘要:
Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language processing (NLP), particularly those based on transformer architectures, have significantly accelerated progress across various NLP tasks. This study focuses on the reliability of transformer-based question answering (QA) models, specifically BERT models and its variants (RoBERTa, ALBERT, DistilBERT). These encoder-only pretrained transformers have demonstrated remarkable accuracy in QA tasks that can be treated as classification tasks. However, their reliability remains underexplored. This study evaluates the reliability of four BERT-based models by assessing response stability under two conditions: (1) internal model variations induced via Monte Carlo Dropout (MCD) and (2) input perturbations through paraphrasing. Using the SQuAD and QuAC datasets, we investigate how dropout rates affect prediction consistency and whether lexical changes impact answer stability. Our findings reveal that RoBERTa maintains higher reliability, whereas AlBERT and DistilBERT exhibit significant inconsistencies. Statistical analyses confirm that enabling MCD during prediction does not disrupt inference dynamics, validating its effectiveness as a reliability metric. These findings underscore the importance of evaluating both accuracy and stability in QA models to ensure stability in real-world applications.