arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13481cs.CL

一种用于生物医学和临床文本抽取式摘要的混合分层1D-CNN-BiLSTM框架

A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text

Saad Bin Ather, Muhammad Saif, Ali Hassan Khan, Manzer Abbas, Hajra Waheed

首次发表
浏览论文内容

中文总结 AI 辅助

提出混合分层1D-CNN-BiLSTM框架,通过抽取式句子选择避免生成式幻觉,在生物医学文本摘要中优于基线,实现事实有据的可靠摘要。

中文摘要 AI 辅助

大型语言模型使得生成式摘要变得异常流畅,但生成的摘要可能产生事实幻觉,这在生物医学和临床领域构成严重风险。我们通过从流程中移除生成环节,将摘要任务重新定义为抽取式句子选择来解决这一问题。我们的混合分层CNN-LSTM摘要器使用堆叠的多核卷积将句子级嵌入组合成更丰富的句子间表示,随后通过双向LSTM对文档中的长距离依赖进行建模。一个轻量级评分头为每个句子分配重要性分数,并以二元交叉熵损失针对oracle抽取标签进行端到端训练。在推理时,采用动态的均值加标准差阈值并辅以top-3回退机制,直接从源文本中选择句子,并按时间顺序重新排列以生成最终摘要。由于每个输出句子都直接复制自输入,该模型避免了由生成引起的事实漂移。在PubMed数据集上,我们的架构优于独立的CNN和LSTM基线,而消融实验表明,更宽的卷积感受野能改善句子评分。在MIMIC-CXR和MIMIC-IV BHC数据集上,该模型在非结构化叙述文本上表现良好,但在高度模板化的报告中则退化为基于位置的基线。这些结果表明,结构约束可以为构建事实有据的摘要系统提供一条可靠路径,这类系统通过设计而非修正来保证可信度。

英文摘要

Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extractive sentence selection. Our Hybrid Hierarchical CNN-LSTM Summarizer uses stacked multi-kernel convolutions to compose sentence-level embeddings into richer inter-sentence representations, followed by a bidirectional LSTM to model long-range dependencies across the document. A lightweight scoring head assigns per-sentence importance scores and is trained end-to-end with binary cross-entropy against oracle extractive labels. At inference, a dynamic mean-plus-standard-deviation threshold with a top-3 fallback selects sentences directly from the source and chronologically reorders them into the final summary. Since every output sentence is copied from the input, the model avoids generation-induced factual drift. On PubMed, our architecture outperforms isolated CNN and LSTM baselines, while ablations show that wider convolutional receptive fields improve sentence scoring. On MIMIC-CXR and MIMIC-IV BHC, the model performs well on unstructured narratives but defaults toward positional baselines on highly templated reports. These results suggest that structural constraints can provide a reliable path toward factually grounded summarization systems that are trustworthy by design rather than by correction.

发表机构

  • NUCES-FAST Lahore(拉合尔国立计算机与新兴科学大学(FAST))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑