发表机构
Doha Institute for Graduate Studies; Asia Pacific University of Technology & Innovation(多哈研究生学院; 亚太科技与创新大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生物医学文本分类中PDF解析引入的噪声,提出保守可审计的拼写纠正可靠性层,弃权(不执行)不确定编辑,可恢复约80.45%的噪声导致的macro-F1下降,保持接近干净的分类性能。
AI 中文摘要
生物医学自然语言处理(NLP)流程通常假设输入文本是干净的,但通过自动PDF解析组装的大规模语料库存在普遍的类OCR人工制品、分词拆分与合并、连字符残留以及字符级损坏,这些问题会系统性削弱词汇证据并降低下游分类器性能。我们提出一种保守、完全可审计的拼写纠正可靠性层,将其设计为面向安全的预处理模块而非追求最高准确率的纠正器:在不确定情况下,系统遵循医疗不伤害原则弃权(不执行)编辑操作。该确定性架构结合了有界编辑距离的候选生成、语料库衍生的n元语法评分以及一套保护领域关键术语的生物医学安全门。我们从内在和外在两方面评估该层:内在评估使用手动整理的包含2104个 token 级案例的基准;外在评估使用跨10000个样本的三分类CORD-19主题分类器(预防、治疗、流行病学),遵循原则性的四运行协议(干净、有噪声、恢复、安全)。内在评估中,该层在合成错误上达到94.61%的错误修复召回率,对负对照零有害编辑。外在方面,它恢复了约80.45%的噪声诱导的macro-F1下降,将macro-F1从(有噪声)的0.7654提升至(恢复)的0.7717,同时保持接近干净的性能(安全:0.7721)。针对用BioBERT分类的103个真实世界OCR提取摘要的补充案例研究证实,Transformer编码器对轻度噪声表现出相对鲁棒性,这推动了未来结合有界神经信号和UMLS词表且不损害可审计性的灰盒架构的开发。该系统完全确定性、人工制品驱动,设计时考虑了部署和可审计性。
英文摘要
Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that systematically erode lexical evidence and degrade downstream classifiers. We introduce a conservative, fully auditable spell-correction reliability layer conceived as a safety-oriented preprocessing module rather than a maximal-accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do-no-harm philosophy. The deterministic architecture couples bounded edit-distance candidate generation with corpus-derived n-gram scoring and a suite of biomedical safety gates that protect domain-critical terminology. We evaluate the layer both intrinsically, on a manually curated benchmark of 2,104 token-level cases, and extrinsically, on a tri-class CORD-19 topic classifier (Prevention, Treatment, Epidemiology) spanning 10,000 examples under a principled four-run protocol (Clean, Noisy, Restored, Safety). Intrinsically, the layer attains 94.61% error-fix recall on synthetic errors with zero harmful edits on negative controls. Downstream, it recovers approximately 80.45% of the noise-induced macro-F1 degradation, elevating macro-F1 from 0.7654 (Noisy) to 0.7717 (Restored) while preserving near-clean performance (Safety: 0.7721). A supplementary case study on 103 real-world OCR-extracted abstracts classified with BioBERT confirms that transformer encoders appeared relatively robust to mild noise, motivating a future grey-box architecture that integrates bounded neural signals and UMLS lexicons without compromising auditability. The system is fully deterministic, artifact-driven, and designed with deployment and auditability in mind.
Comments6 pages. Accepted at the 2026 14th International Conference on Bioinformatics and Computational Biology (ICBCB 2026), Kitakyushu, Japan; to appear in IEEE Conference Proceedings (IEEE Xplore). Code: https://github.com/MoustafaMohamedMoustafaHassan/the_signal_in_the_noise_CORD19_project