arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

零样本医学文摘分类中DeBERTa-v3的比较可解释性框架

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

arXiv 2610.02116首次发表:更新:

发表机构

Universidad de Castilla-La Mancha; University of Murcia; Murcian Bio-Health Institute (IMIB-Arrixaca)(卡斯蒂利亚-拉曼恰大学; 穆尔西亚大学; 穆尔西亚生物健康研究所(IMIB-Arrixaca))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出比较可解释性框架审计DeBERTa-v3零样本医学文摘分类,通过五种归因方法比较和Jaccard一致性量化,揭示解释稳定性与预测确定性正相关,并识别三种系统性失败机制,支持多方法联合审计。

AI 中文摘要

本文提出了一个比较可解释性框架,用于审计DeBERTa-v3在医学文摘零样本分类中的表现。该工作解决了可解释人工智能中的分歧问题,即不同的归因方法对相同的输入和预测产生不同的解释。在医学文摘语料库上实现了一个自然语言推理引擎,每个诊断类别使用五个增强假设,每类使用一千个文本的平衡样本。比较了五种解释方法:SHAP和LIME作为模型无关方法,遮挡和输入×梯度作为深度学习特定方法,以及注意力×梯度作为Transformer特定方法。解释通过top-token归因进行标准化,并使用Jaccard指数量化成对一致性。在定义明确的临床领域实现了高预测准确率,而在高语义模糊性下性能下降。解释稳定性直接反映预测确定性,在单值类别中表现出强收敛,在诊断不确定性下显著下降。此外,定性错误审计揭示了三种系统性失败机制:词汇过敏、语义重叠和归因连贯性丧失。结果支持在审计医学文本分类中的Transformer模型时,结合使用多种解释方法和定量一致性指标,并建议优先考虑特定的临床本体论而非宽泛的诊断标签。

英文摘要

A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

Comments18 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑