AI 中文总结
本研究评估LLMs短、长上下文窗口生成的文献综述,发现长上下文LLMs虽能纳入更多信息但存在重复等问题,AI生成综述需人工完善,未来应采用人机结合混合方法解决局限。
AI 中文摘要
本研究聚焦于评估大语言模型(LLMs)在短上下文与长上下文设置下生成的文献综述,以探究上下文窗口对AI生成文献综述质量的影响,以及AI在辅助文献综述写作中的作用。两名研究人员针对基于Semantic Scholar和Arxiv的研究源生成的20篇AI文献综述,从15个维度开展评估。研究发现,AI生成的文献综述需人工监督才能达到学术出版标准;随着上下文窗口增大,LLMs可纳入更广泛信息并在更长输入中保持连贯性,但也加剧了内容重复、关键研究遗漏、偏向描述性而非综合性等问题。本研究表明,AI生成的综述可提供基础概述,但其输出必须由领域专家严格评估与完善;未来研究应考虑整合其他LLMs及微调模型,采用结合人类专业知识与AI能力的混合方法,以解决本研究识别出的局限性。
英文摘要
Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.