ThaiTrees:跨领域的泰语句法依存树
ThaiTrees: Thai Syntactic Dependency Trees Across Domains
- Chulalongkorn University(朱拉隆功大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对泰语缺乏大规模自动解析语料库的问题,构建了包含3.42亿词符、覆盖多领域的ThaiTrees语料库,并开发了可复现的清洗、处理与解析流程,支持句法分布研究。\n
AI中文摘要:
研究自然语言中的句法模式需要大规模已解析语料库,但人工标注成本高昂且难以扩展。泰语拥有用于训练和评估解析器的人工标注依存树库,但缺乏用于定量句法研究的大规模自动解析语料库。我们提出了ThaiTrees,一个包含3.42亿词符的语料库,数据来源于新闻、维基百科、口语转录文本和社交媒体。我们开发了一个可复现的流程,用于在通用依存框架下对泰语文本进行清洗、处理和解析。生成的语料库使语法关系可检索,并支持句法分布研究。我们发布了频率词典和CoNLL-U格式的解析结果,采用适合AI辅助和传统程序化分析的机器可读格式。
英文摘要:
Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.