JEV 与 LLM:七项政治学复现研究中的准确性、成本与校准
JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
浏览论文内容
中文总结 AI 辅助
本文通过七项政治学复现任务,对比 JEV 与 LLM 及人类编码者的准确性、成本和校准,发现 JEV 性能接近 LLM,但无成本优势,主要优势在于解析选择概率的便利性。
中文摘要 AI 辅助
大型语言模型(LLM)通过生成文本标记来标注和量化政治文本或建构。一类新型模型(TypeSafe 将其作为“System One”模型销售)则返回用户提供的固定答案集上的决策和概率分布。商业模型 JEV 被宣传为相比传统 LLM 具有显著的成本和速度优势,同时决策校准更佳。因此,对于希望快速且经济地标注或量化大规模文本语料库,并需要分类器不确定性的可靠指标的社会科学家而言,JEV 可能很有用。然而,这些声称的准确性以及该模型在社会科学文本任务中的更广泛准确性尚未得到证实。在本文中,我们正是要做这项工作,并希望确立 JEV 对社会科任务的适用性。我们将 JEV 与已发表研究中的 LLM 和人类编码者进行比较,并与当前中端商业 LLM(GPT-6 Luna)和开放权重替代模型(Qwen3.8-27B)进行比较。我们发现,JEV 在各种任务中与这两种 LLM 的能力相当或接近。然而,我们发现,在 OpenAI 的批量价格下,JEV 相比 GPT-6 Luna 没有成本优势。此外,我们发现,当每个问题只询问一次时,JEV 的概率校准优于 GPT-6 Luna 的标记概率,但并不始终优于 Qwen3.8-27B。我们得出结论,除非研究人员有速度需求,否则 JEV 唯一明显的优势是易于解析底层选择概率。
英文摘要
Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.
发表机构
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。