arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NLP驱动的古印度翻译医学文本知识提取与主题分类

NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

M. S. Rajeevan, B. Mini Devi, V. S. Anoop, C. Mallikarjuna

arXiv 2608.28608首次发表:更新:

发表机构

University of Kerala; Thiagarajar School of Management (Autonomous); Indian Institute of Technology Hyderabad (IITH)(喀拉拉大学; 蒂亚加拉贾管理学院(自治); 印度理工学院海得拉巴分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究采用NLP相关技术对古印度翻译医学文本进行知识提取与主题分类,实现阿育吠陀医学智慧的计算组织,提升其可获取性,推动多学科发展。

AI 中文摘要

古印度医学文本如《妙闻集》(Sushruta Samhita)包含大量关于疾病、治疗方法和外科技术的信息,但其古老的格式与复杂的词汇给获取和系统整理带来困难。本研究利用自然语言处理(NLP)方法,包括命名实体识别(NER)、BERTopic建模和Neo4j知识图谱开发,基于翻译版本对重要概念进行提取、分类与可视化。使用BERTopic的主题分类可识别潜在的医学主题,NER支持对疾病、治疗方法、研究者和药用植物的结构化实体识别;通过Neo4j的基于图的网络分析还可对提取实体间的关系进行语义表示,支持知识检索与数字保存。研究结果表明,图数据库、主题建模和实体识别有助于对阿育吠陀历史医学智慧进行计算组织,缩小传统文本与当代数据驱动研究之间的差距。所提方法推动了历史文本分析、医学信息学与数字人文学科的发展,使古印度医学智慧更易获取和理解。

英文摘要

Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions. Thematic classification with BERTopic allows for the identification of the underlying medical topics, whereas NER supports the structured entity recognition of diseases, treatments, researchers, and medicinal plants. Graphbased network analysis with Neo4j also allows for the semantic representation of relationship among extracted entities, supporting knowledge retrieval and digital preservation. The findings illustrate how graph databases, topic modeling, and entity recognition facilitate the computational organization of Ayurveda's historical medical wisdom, closing the gap between the conventional texts and contemporary data-driven inquiry. The suggested method promotes historical text analysis, medical informatics, and digital humanities to make ancient Indian medical wisdom more accessible and understandable.

Comments19 pages, 8 figures, 5 tables. Presented at the National Conference on "Reimagining LIS Education: Integrating Indian Knowledge Systems with NEP 2020" (March 2025), organized by Tata Institute of Social Sciences (TISS) and the Indian Association of Teachers of Library and Information Science (IATLIS). Recipient of the Best Paper Award

Journal refIn Proc. TISS-IATLIS National Conference 2025: Reimagining LIS Education: Integrating Indian Knowledge Systems with NEP 2020, Vol. 1, p. 351, 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑