发表机构
Department of Computer Science, University of Delhi; Department of Statistics, Ramjas College, University of Delhi(德里大学计算机科学系; 德里大学拉姆贾斯学院统计系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型零样本摘要的稳定性与可信度问题,提出两级诊断协议,通过文档级稳定性分析及分层样本观察估计稳定性指数,经对三种文档类型的三个LLM摘要器实证研究,揭示差异,推动相关研究发展。
AI 中文摘要
使用大语言模型(LLMs)的零样本摘要通过生成连贯流畅的摘要显著推进了抽象摘要任务。然而,大语言模型的潜在随机性引发了对LLM生成摘要的稳定性和可信度的担忧。由于LLM生成的摘要在教育环境中激增,这个问题变得越来越重要。我们提出了一种新颖的两级诊断协议,用于基于生成摘要的稳定性对LLM摘要器进行基准测试。在较低级别,在受控环境下生成的多个LLM摘要上进行文档级稳定性分析,并计算稳定性系数。对每个生成的摘要与原始文档的语义和事实一致性进行评分,从而能够在多个维度上估计稳定性。在下一级别,对从语料库中抽取的分层文档样本的观察结果进行整合,以估计LLM摘要器的稳定性指数,该指数是其可信度的代理。我们对三种文档类型的三个LLM摘要器的实证研究揭示了不同LLM在摘要评估指标上的生成级变异性存在统计学上的显著差异。这项研究通过对LLM摘要中稳定性问题的证据识别推进了LLM摘要研究,并推动了对强大、可靠和可信的LLM摘要器开发的进一步研究。
英文摘要
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries. However, underlying stochasticity of the large language models raises concerns about the stability and trustworthiness of the LLM-generated summaries. This issue has become increasingly important due to proliferation of LLM-generated summaries in educational settings, where students and researchers summarize complex academic materials in zero-shot manner. We propose a novel two-level diagnostic protocol for benchmarking LLM-summarizers based on the stability of the generated summaries. At the lower level, document-level stability analysis is performed over multiple LLM-summaries generated under controlled environment, and the stability coefficient is computed. Each generated summary is scored for semantic and factual alignment with the original document, enabling estimation of stability along more than one dimensions. At the next level, observations from a stratified sample of documents drawn from the corpus are consolidated to estimate the stability index of the LLM-summarizer, which is the proxy for its trustworthiness. Our empirical investigation of three LLM-summarizers across three genres of documents reveals statistically significant differences in the generation-level variability among LLMs across summary evaluation metrics. This study advances the LLM-summarization research by evidential recognition of the stability problem in LLM-summaries and motivates further research towards development of robust, reliable and trustworthy LLM-summarizers.
Comments28 pages, Under review in a journal