AI 中文总结
研究探讨生成式人工智能模型中概率性“复制”问题,指出能从大语言模型提取版权作品文本,因其权重存储令牌统计关系影响概率性生成过程。版权法对此判定重要但现有法律难助力,认为可能采取功能性方法,还建议了法律变革。
AI 中文摘要
近期研究表明,能从一些大语言模型中提取某些版权作品的逐字或近乎逐字的文本,这证明模型权重以某种形式编码了这些作品,即模型从训练数据中“记住”了它们。但大语言模型存储信息的格式与常见数据库不同,其权重存储从训练数据中学到的令牌之间的统计关系,这种关系影响的生成过程通常是概率性而非确定性的。在记忆的情况下,这些关系足够强,以至于在许多情况下,模型可能会以一定概率从其训练数据中生成版权作品。版权法此前无需判定存储可能产生或不产生与版权作品相似输出的信息本身是否是该作品的复制。这个问题的答案很重要,因为它可能决定许多大语言模型的合法性。成文法和判例法对此帮助不大。我们认为版权法可能会采取功能性方法来处理这个问题,即只有当能直接从输出中提取特定作品时,才认定大语言模型包含该作品的复制。作为政策问题,这一结果并不令人满意,我们建议了法律的潜在变革,但这是现行法律下最可能的结果。
英文摘要
Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability. Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work. The answer to the question is important, because it may determine the legality of many LLMs. The statute and case law are largely unhelpful. We argue that copyright law will likely take a functional approach to the question, finding that LLMs contain a copy of a particular work only if it is straightforward to extract that work in outputs. That result is unsatisfying as a policy matter, and we suggest potential changes to the law, but it is the most likely outcome under current law.
CommentsForthcoming, Berkeley Technology Law Journal