发表机构
University of Pisa; University of Trento; RazorPay(比萨大学; 特伦托大学; RazorPay)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个探究LLMs理解纵向政治贸易谈判潜力的基准TradeVerse,基于WTO贸易关切构建,含三项任务,实验凸显当前LLMs面临的挑战。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被应用于涉及制度和政治文本的任务,但现有基准多在孤立文档或单一任务上对其进行评估。在现实政治中,谈判属于纵向数据,参与方会在多轮中达成一致或展开争论,每一轮都是前序轮次的结果,因此理解某一轮需追踪此前所有内容。本文介绍TradeVerse——一个基于世界贸易组织(WTO)特定贸易关切构建的基准,其中成员国会相互提出质疑并在多轮(有时长达数年)中交换论点。在TradeVerse中,我们重构了1170场会议的记录,涵盖5个组别和89个产品组别,并定义了三项任务:第一,系统需分析纵向会议记录,预测特定会议中所讨论产品的协调制度编码(HS章节);第二,我们检验系统在分析匿名化会议内容后,能否猜出回应国的名称;第三,我们要求系统扮演回应国,为最后一轮提供发言。所有标签均直接从会议记录中获取,无需人工标注。我们的实验凸显了这些任务对当前LLMs构成的挑战。据我们所知,TradeVerse是首个探究LLMs在理解纵向政治贸易谈判方面潜力的基准。
英文摘要
LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating parties can align or argue over multiple iterations and each turn is an outcome of the previous turns, hence, understanding one turn requires tracking everything before it. We introduce TradeVerse, a benchmark built from the World Trade Organisation (WTO) specific trade concerns, where member states challenge one another and exchange arguments over multiple rounds, sometimes for years. We, in TradeVerse, reconstruct minutes of $1170$ meetings, spanning across 5 groups and $89$ product groups and define three tasks: first, the system has to analyze the longitudinal meeting records and predict the harmonized system codes (HS chapters) of the products under discussion in the particular meeting, second, we examine whether the system, upon analyzing the anonymized content of the meeting, can guess the name of the responding country and third, we ask the system to play the role of the responding country and provide the statement for the very last round. All labels are recovered directly from the proceedings, requiring no manual annotation. Our experiments highlight the challenges these tasks pose for current LLMs. To the best of our knowledge, TradeVerseis the first benchmark to investigate potential of LLMs in understanding longitudinal political trade negotiations.