医学教育中针对大型非结构化文本数据的检索增强生成与代表性向量摘要
Retrieval Augmented Generation and Representative Vector Summarization for large unstructured textual data in Medical Education
- University of Peradeniya(佩拉德尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文探讨了检索增强生成在医学教育领域的应用,并提出了一种基于代表性向量、结合抽取式与抽象式的大型非结构化文本数据摘要方法。
AI中文摘要:
大型语言模型正日益广泛地用于包括内容生成和聊天机器人在内的各种任务。尽管在通用任务中表现出色,但在应用于特定领域任务时,LLM需要进行对齐,以缓解幻觉和产生有害答案的问题。检索增强生成(RAG)允许轻松地将非参数化知识库附加到LLM并进行操作。本文讨论了RAG在医学教育领域的应用。提出了一种使用代表性向量针对大型非结构化文本数据的抽取式与抽象式相结合的摘要方法。
英文摘要:
Large Language Models are increasingly being used for various tasks including content generation and as chatbots. Despite their impressive performances in general tasks, LLMs need to be aligned when applying for domain specific tasks to mitigate the problems of hallucination and producing harmful answers. Retrieval Augmented Generation (RAG) allows to easily attach and manipulate a non-parametric knowledgebases to LLMs. Applications of RAG in the field of medical education are discussed in this paper. A combined extractive and abstractive summarization method for large unstructured textual data using representative vectors is proposed.