(谁的默认设置?)人工智能是否在重新定向考古学方法?
(Whose defaults?) Is artificial intelligence reorienting archaeological methods?
- Seminar für Ur- und Frühgeschichte(史前与早期历史研讨班)
- Georg-August-Universität Göttingen(哥廷根大学)
- McDonald Institute for Archaeological Research(麦克唐纳考古研究所)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过分析119,000篇摘要和受控实验,发现大语言模型虽未减少考古学实际方法多样性,但其推荐趋于收敛,引发对方法多样性保持的思考。
AI中文摘要:
生成式人工智能和“氛围编码”实践正在改变考古学家开展计算研究的方式,但它们对该学科方法范围的影响仍未得到充分研究。在本文中,我们评估大语言模型(LLMs)是否正在缩小考古学家使用的方法多样性。我们首先分析了来自Scopus的约119,000篇考古学摘要,涵盖2010年至2025年的出版物。使用本地运行的LLM,我们识别了每篇摘要中报告的计算方法,并将其组织为25个广泛类别(L2)和241个更细的聚类(L3)。一个关于子学科内方法组成的贝叶斯Dirichlet-多项模型发现,2023年之后方法使用出现了微小但可信的转变。然而,这一转变小于整个研究期间已经存在的变异。没有单一技术显示出显著变化,总体方法多样性反而增加了。随后,我们进行了一项受控实验,以检验LLMs是否推荐比考古学家实际使用过的方法范围更窄的方法集。两个不同的开放权重模型被要求为28个考古学研究问题建议方法,提示提供了三个级别的方法指导:新手、中级和专家。推荐多样性远低于已发表文献中的多样性,尤其是在没有方法指导的情况下。模型还倾向于偏爱2023年之前广泛使用的方法,且其推荐更接近2023年后的文献。综合来看,这些结果与LLMs推动方法选择趋于收敛一致,尽管我们的研究无法确立因果效应。它们提出了一个更广泛的问题:随着LLMs更多地参与研究,考古学如何保持方法多样性?
英文摘要:
Generative AI and the practice of "vibe coding" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?