AI 中文总结
研究LLM智能体版权法合规性评估问题,通过引入Copyright - Bench基准,由网站开发等商业任务组成,含提示变化与时间压力,对比智能体与人类基线,发现智能体存在选择版权作品及违规率受偏好和压力影响等情况。
AI 中文摘要
大语言模型(LLM)智能体越来越多地执行涉及检索外部内容(如图像)并在适当情况下复制该内容的商业任务。LLM智能体应遵守包括版权法在内的法律,但目前缺乏评估它们在实际中是否合规的适当框架。为此,我们引入了Copyright - Bench,这是一个旨在评估LLM智能体对版权法合规性的基准。它由现实商业任务组成,包括网站开发、商品设计和推销资料制作等,涉及智能体在公共领域内容(使用合法)和版权内容(在此设置下使用侵权)之间进行选择。评估引入了模拟不同用户偏好的提示变化以及时间压力。将最先进的LLM智能体与人类基线进行比较,我们发现:(1)即使有公共领域替代方案,智能体仍会选择版权作品;(2)对于开放权重模型,违规率会因某些用户偏好和模拟时间压力而增加。
英文摘要
Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content, such as images, and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce Copyright-Bench, a benchmark designed to evaluate LLM agents' compliance with copyright law. Copyright-Bench comprises realistic commercial tasks---website development, merchandise design, and pitch deck production---that involve agents selecting between public-domain content, the use of which is legal, and copyrighted content, the use of which is infringing in this setting. The evaluation introduces prompt variations that simulate different user preferences, as well as time pressure. Comparing state-of-the-art LLM agents against a human baseline, we find that: (1) agents select copyrighted works despite the availability of public-domain alternatives; and (2) for open-weight models, violation rates increase in response to certain user preferences and simulated time pressure.
CommentsICML 2026 Spotlight