LLM与基于规则的标注在推理叙事特征上的评分者间信度:土耳其语语料库的三项研究
Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
浏览论文内容
中文总结 AI 辅助
本研究通过三项实验评估了LLM与规则检测器在土耳其语语料库推理叙事特征标注上的信度,发现机器与人类一致性极低,表明该特征难以自动检测或规则定义不够可操作。
中文摘要 AI 辅助
带有自动生成特征标注的数据集引发了一个很少被问及的问题:人类会同意这些标签吗?本报告针对Objective Projection语料库回答了这一问题,该语料库是一个土耳其语叙事数据集,其场景带有基于规则的检测器针对六种创作手法特征生成的每场景applied_rules字段——包括两种禁止性特征(情感标注、明喻)和四种积极性手法(具象隐喻、微聚焦、时间锚定、氛围矛盾)。报告包含三项研究。研究1(n=120)将检测器与方案作者本人的盲标签进行对比评分。研究2(n=100,为不相交的场景集)将检测器以及Gemini 2.5 Flash和Grok与一位独立非专家评分者的结果进行对比,该评分者的标签在任何机器运行前已锁定。研究2b使用Claude Fable 5(High)和ChatGPT 5.5重复了完全相同的协议。核心结果涉及一条规则。在具象隐喻上——该方法论最接近理论核心的特征——五个机器标注器在100个场景中分别返回了0、1、40、72和78个阳性标签,而人类计数为9。Cohen's κ值对于六个标注器中的五个,在两个人参考标准和两个场景集上均处于或无法与随机水平区分:0.004、0.015、0.000、0.019、0.027。原始一致性范围从74.7%到84.5%,这是类别不平衡的产物,而非能力的标志。我们刻意不将其归结为单一结论。两种解读并存:该特征确实具有推理性质,超出了当前自动检测的能力;或者该规则的定义尚不足以让任何评分者(包括人类)一致应用。区分这两种解读需要第二位独立的人类评分者,而本报告没有,因此也不做此声称。
英文摘要
Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field from a rule-based detector over six craft features -- two prohibitions (emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Three studies are reported. Study 1 ($n = 120$) scores the detector against blind labels from the scheme's own author. Study 2 ($n = 100$, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5 (High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor -- closest to the methodology's theoretical core -- the five machine labellers returned positive rates of $0$, $1$, $40$, $72$ and $78$ out of $100$ scenes, against a human count of $9$. Cohen's $κ$ was at or indistinguishable from chance for five of six labellers, across both human references and both scene sets: $0.004$, $0.015$, $0.000$, $0.019$, $0.027$. Raw agreement ranged from $74.7\%$ to $84.5\%$, an artefact of class imbalance rather than a sign of competence. We deliberately do not resolve this into a single story. Two readings survive: the feature is genuinely inferential and beyond current automatic detection, or the rule's definition is not yet operational enough for any rater to apply consistently -- including the human. Distinguishing them needs a second independent human rater, which this report does not have and therefore does not claim.
发表机构
- Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。