arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非结构化生物医学文本标注的一致性度量

Consensus Measures for Unstructured Biomedical Text Annotations

Pascal Wullschleger, Christian Kreis, Martin A. Walter, Jennifer Foster, Marc Pouly

arXiv 2608.03529首次发表:更新:

AI 中文总结

该研究针对生物医学非结构化文本标注的一致性量化难题,提出基于自然语言推理的折中度量方案,经实验验证多种语义等价度量的特性及局限性。

AI 中文摘要

生物医学文献正被越来越多地用于挖掘其原始创作目的之外的知识。由于目标概念无法预先确定,标注者更倾向于开放式标签,而这类标签的一致性难以量化。我们针对生物医学标注任务中提供非结构化文本的标注者,研究软交互标注者可靠性。合成实验表明,软可靠性可通过多种语义等价度量来量化,且度量的选择会影响估计的失效模式;嵌入方法可扩展,但在区分相似却不同的概念时存在局限;大语言模型颇具潜力,但在估算偶然一致性时受限于可扩展性。最后,我们提出基于自然语言推理的度量方案,作为合理的折中选择。

英文摘要

Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard to quantify. We study soft inter-rater reliability for annotators providing unstructured texts for biomedical annotation tasks. Synthetic experiments show that soft reliability can be quantified using a variety of semantic equivalence measures, and that the choice of measure affects failure modes of the estimation. Embeddings are scalable, but limited when differentiating similar but distinct concepts. Large language models are promising, but limited by scalability for estimating agreement by chance. Finally, we suggest measures based on natural language inference as a sensible compromise.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑