Viveka-Insight:基于引文检索的跨语言概念图与检索资源,覆盖斯瓦米·维韦卡南达英文与孟加拉文全集
Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali
浏览论文内容
中文总结 AI 辅助
Viveka-Insight构建了维韦卡南达英孟双语非平行语料库的跨语言概念图与检索资源,通过四层结构实现引文可验证的跨语言检索,精确率0.60,可迁移至其他经典语料库。
中文摘要 AI 辅助
经典哲学语料库对语言资源提出了三个叠加的挑战:它们以多种语言存在且缺乏平行对齐,其词汇远离当代读者的语言,并且对文化敏感材料的生成文本必须可验证地基于引文。我们提出了Viveka-Insight,一个针对斯瓦米·维韦卡南达(1863-1902)作品的双语资源和开源流水线:九卷英文全集和十卷孟加拉文Vani o Rachana,这两个相关但非平行的语料库约含1500万字符。我们发布了四个层次:(i)保留结构的解析(32,694个段落,168,842个句子),每个段落带有深度链接到出版版本的锚点;(ii)一个包含8,362个语言无关概念的跨语言概念图,具有87,518条类型化的段落-概念边和55,872条概念-概念边,其中规范的英文标签作为字符串相等键,在没有平行数据的情况下连接孟加拉文和英文段落;(iii)一个包含60,850个表面形式(30,053个英文,30,797个孟加拉文)的双语别名清单;(iv)一组由三位标注者评判的200条人工标注的段落-概念边,并发布所有逐标注者的评判。我们报告了在194个经过验证的演讲对上的已知项跨语言检索(两个方向上的Recall@10均为0.86),一项包含30个问题的引文完整性和现代问题桥接审计,以及一项在严格的两标注者共识下概念提取精确率为0.60(Cohen's kappa = 0.61)的人类研究。提取器的置信度权重经过校准:限制权重>=0.8可将精确率提高到0.71,同时保留98%的含概念段落。孟加拉文的精确率明显低于英文(0.54对0.68),将弱点定位在跨语言访问所依赖的那一半。该设计可迁移到其他多语言经典语料库。
英文摘要
Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight, a bilingual resource and open-source pipeline for the works of Swami Vivekananda (1863-1902): the nine-volume English Complete Works and the ten-volume Bengali Vani o Rachana, two related but non-parallel corpora of about 15 million characters. Four layers are released: (i) a structure-preserving parse (32,694 paragraphs, 168,842 sentences) with per-paragraph anchors deep-linking into the published editions; (ii) a cross-lingual concept graph of 8,362 language-agnostic concepts with 87,518 relation-typed paragraph-concept and 55,872 concept-concept edges, in which canonical English labels act as a string-equality key linking Bengali and English passages with no parallel data; (iii) a bilingual alias inventory of 60,850 surface forms (30,053 English, 30,797 Bengali); and (iv) a human-annotated set of 200 paragraph-concept edges judged by three annotators, released with all per-annotator judgments. We report known-item cross-lingual retrieval over 194 verified rendered lecture pairs (Recall@10 0.86 in both directions), a 30-question audit of citation integrity and modern-question bridging, and a human study placing concept-extraction precision at 0.60 under strict two-annotator consensus (Cohen's kappa = 0.61). The extractor's confidence weight is calibrated: restricting to weight >= 0.8 raises precision to 0.71 while retaining 98% of concept-bearing paragraphs. Precision is markedly lower in Bengali than English (0.54 vs 0.68), locating the weakness in exactly the half that cross-lingual access depends on. The design transfers to other multilingual classical corpora.
发表机构
- Ramakrishna Mission Vivekananda Educational and Research Institute(罗摩克里希那传教会维韦卡南达教育与研究学院)
机构由 AI 辅助整理,请以论文原文为准。