大语言模型的数据引用:一项挑战
Data Citation for Large Language Models: A Challenge
浏览论文内容
中文总结 AI 辅助
本文指出大语言模型的数据引用是一项开放挑战,提出训练数据归因、推理时数据引用、知识图谱事实引用三个研究方向,需跨领域协同推进以解决该问题。
中文摘要 AI 辅助
大语言模型越来越多地成为人们获取信息的媒介,越来越多的研究关注它们是否会引用其输出背后的来源。这些研究将引用视为一种验证手段,并将其应用于文本文档。学术引用还有另外两个功能:署名和溯源,且它对数据和文本同样适用。本文认为,大语言模型的数据引用是一项开放的挑战,它与文档级引用定位不同,且更难解决。我们提出,这类模型应如何引用数据,才能使输出保持可验证性、溯源保持可追踪性,且署名能传递给数据创建者和整理者。我们提出了三个研究方向:训练数据归因必须将影响力估计转化为对被吸收进模型参数的语料库的引用;推理时的数据引用必须以合适的粒度和固定性识别数据集、子集和查询结果;引用知识图谱事实必须定义对单个三元组的引用表示什么,以及署名如何沿溯源传递。这三个方向的进展都依赖于数据库、信息检索、知识表示和人工智能领域的协同工作。
英文摘要
Large language models increasingly mediate access to information, and a growing body of work asks whether they cite the sources behind their outputs. That work treats citation as a verification device and applies it to textual documents. Scholarly citation serves two further functions, credit and provenance, and it applies to data as much as to text. This paper argues that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve. We ask how such models should cite data so that outputs stay verifiable, provenance stays traceable, and credit reaches data creators and curators. We set out three research directions. Training data attribution has to turn influence estimates into references for corpora absorbed into model parameters. Data citation at inference time has to identify datasets, subsets, and query results at the right granularity and fixity. Citing knowledge graph facts has to define what a reference to a single triple denotes and how credit propagates along provenance. Progress on all three depends on joint work across the database, information retrieval, knowledge representation, and artificial intelligence communities.
发表机构
- University of Padua(帕多瓦大学)
机构由 AI 辅助整理,请以论文原文为准。