JieZi:用于古文字训诂的大规模专家审核数据集与基准
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
浏览论文内容
中文总结 AI 辅助
本文针对古文字训诂研究缺乏结构化数据与基准的问题,提出ACCE任务,构建JieZi-Dataset与JieZi-Bench,实验发现多模态大模型在古文字训诂的高阶任务上表现不佳,微调后性能显著提升。
中文摘要 AI 辅助
古文字的学术训诂需要整合视觉观察、语言分析与历史语境。然而,现有计算方法仅聚焦于文字识别、检索等子任务,缺乏开展全面学术分析所需的结构化数据集与基准。为解决这一局限,本文提出古文字训诂(Ancient Chinese Character Exegesis, ACCE),这是一项建模学术训诂过程的视觉-语言问答(Vision-Language Question Answering, VQA)任务。ACCE分为四个递进层级:基础文字识别、字形分析、意义训诂与历时演变分析。为支撑该任务,本文构建了两类互补资源:JieZi-Dataset是首个面向ACCE的大规模、经专家审核的VQA训练数据集,包含超过50万组问答对。该数据集通过一套流水线构建,通过专家设计的模板与源文本参考约束生成过程以减少事实错误,且在每个关键阶段均进行人工验证以确保学术准确性。JieZi-Bench是与训诂过程对齐的评估基准,由人工专家构建并验证以确保评估可靠性,包含四个层级,其参考答案选自权威辞书且与训练数据分离。对多模态大语言模型的实验表明,当前模型在基础识别任务上表现良好,但在字形分析、语义推理与历时理解方面存在困难;在JieZi-Dataset上进行微调可显著提升四个层级的性能。代码与数据集可在该https URL获取。
英文摘要
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi-Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.
发表机构
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。