arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24574cs.CLcs.CY

评估计算社会科学中文本标注的决策模型

Evaluating Decision Models for Text Annotation in Computational Social Science

Hazem Ibrahim, Yasir Zaki

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估决策模型在18个计算社会科学文本标注任务上的表现,发现其准确率略逊于LLM但成本大幅降低,且置信度校准良好,可作为标注流程中低成本预筛选步骤。

中文摘要 AI 辅助

计算社会科学越来越依赖大型语言模型进行文本标注,已发表研究结果的有效性现在取决于此类模型生成的标签。决策模型是一种为分类问题回答而构建的新模型类别,它以选择、标签集上的概率分布和置信度分数来回答类型化问题,而非自由文本,其推理价格仅为前沿模型的一小部分。这些模型的答案是否准确,以及其所述置信度在社会科学构念上是否可信,尚属未知。在此,我们参照Ziems等人(2024)的评估方法,在18个计算社会科学分类任务(共7,977个条目)上,将首个商业决策模型和两个开放权重对应模型与19个前沿及开放权重语言模型在相同的零样本协议下进行比较。决策模型在15个评估任务中的14个上落后于每个任务的最佳LLM,中位数差距为11.6个宏F1点,而测量成本中位数低44倍。其置信度校准优于19个LLM中16个的言语化置信度,但三个前沿模型显示出更低的校准误差中位数(0.157对0.066)。虽然置信度高于0.9的条目通常被准确标注(中位准确率为0.815),但在一个任务(同伴支持对话中的共情)上,模型报告高置信度却表现接近随机。尽管如此,我们的结果表明,决策模型作为标注流程的第一步是有用的:将低置信度条目路由到LLM,其性能匹配或超过单独使用LLM,而成本仅为后者的四分之一到一半。

英文摘要

Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models. Decision models, a new model class built for categorical question answering, answer typed questions with a choice, a probability distribution over the label set, and a confidence score rather than free text, at a small fraction of frontier inference prices. Whether their answers are accurate, and whether that stated confidence can be trusted on social science constructs, are unknown. Here, we mirror the evaluation of Ziems et al. (2024) on 18 computational social science classification tasks (7,977 items), comparing the first commercial decision model and two open-weight counterparts against 19 frontier and open-weight language models under the same zero-shot protocol, and extending the decision-model comparison to eleven open-weight systems released in the week after it. The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost. Its confidence is better calibrated than the verbalized confidence of 16 of the 19 LLMs, yet three frontier models show lower median calibration error (0.157 against 0.066). While items above 0.9 confidence are typically labeled accurately (median accuracy 0.815), on one task, empathy in peer-support dialogues, the model reports high confidence while performing near chance. Nonetheless, our results suggest that decision models are useful as a first step in the annotation pipeline: routing low-confidence items to an LLM matches or exceeds the LLM alone at a quarter to half of its cost.

发表机构

  • New York University Abu Dhabi(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑