发表机构
The Hong Kong Polytechnic University; Peking University(香港理工大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出保留夹注上下文依赖的计算框架,将集注编纂化为NLP任务,在《山海经》案例中CoNLL F1超97%,为历史注疏知识大规模组织及下游任务奠定基础。
AI 中文摘要
夹注与集注是儒家注疏传统中重要的学术交流形式,但在计算领域鲜有研究关注。本文结合中国传统注疏学与文献学,将集注编纂形式化为一项自然语言处理任务,提出一种计算框架,在实现夹注自动编纂与注疏知识组织的同时保留其上下文依赖关系。该框架采用两步提示链识别注释对应的正文段落与注疏功能,并结合跨源提及聚类整合不同版本的注疏;在《山海经》案例研究中,其CoNLL F1值超过97%。该框架为大规模组织历史注疏知识奠定基础,进而支持广泛的下游文献学与自然语言处理任务。
英文摘要
Inline notes and collected commentaries are important forms of scholarly communication that evolved within the Confucian exegetical tradition, yet have received little computational attention. Drawing on traditional Chinese exegetics and philology, this paper formulates collected commentary compilation as an NLP task and proposes a computational framework that preserves the contextual dependency of inline notes while enabling their automatic compilation and exegetical knowledge organization. It combines two-step prompt chaining for identifying the associated main-text segments and exegetical functions of annotations with cross-source mention clustering for integrating commentary across editions, achieving a CoNLL F1 score above 97% in a case study on the Classic of Mountains. Our framework lays the foundation for the large-scale organization of historical exegetical knowledge, thereby supporting a broad range of downstream philological and NLP tasks.
Comments15 pages, 4 figures