评估自然语言中的理性缔约
Evaluating Rational Contracting in Natural Language
浏览论文内容
中文总结 AI 辅助
本研究针对自然语言缔约的评估缺口,构建理性框架并开发ContractSim评估套件,发现当前LLM智能体在高不确定性下缔约及执行协议存在不足,为相关智能体设计指明改进方向。
中文摘要 AI 辅助
基于语言的AI智能体的出现有望改变机器经济活动的范围,这类智能体不再仅能提出报价或遵循硬编码协议,还可用于以开放式自然语言协商和执行协议。然而,对这些能力的多数评估聚焦于一次性交互或简单经济博弈,未涉及语言可表达的、包含时间延伸、或有条件及不完整的丰富协议空间;同时这些评估侧重原始利润,未衡量可信缔约所需的品质。针对此,我们构建了一个理性框架,用于指导智能体在不确定的多步骤环境中协商和执行自然语言协议。在该框架内,我们开发了量化理性与合作博弈的指标及基线。为评估智能体在此类缔约中的表现,我们将框架实例化为ContractSim(一个评估套件),其中两名智能体在环境及智能体间不确定性下协商并执行多轮供应商协议。在6种环境和3种供应商场景(餐饮、酒店清洁、AI托管)中,我们发现当前基于LLM的智能体能可靠达成协议,且在环境不确定性低时可协商出高效协议;但在高不确定性下,它们常无法协商出可满足、高效或互利的协议,且执行协议时频繁不合作,即便协议易满足也会违反条款以获取额外利润。这些发现凸显了在设计能理性且合作地协商、解释和执行协议的语言智能体方面仍有改进空间。
英文摘要
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in open-ended natural language. However, most evaluations of these abilities have focused on one-off exchanges or simple economic games, leaving open the rich space of time-extended, contingent, and incomplete contracts made expressible by language; they also focus on raw profit, without measuring the qualities required for trustworthy contracting. We address this by formulating a rational framework for how agents should negotiate and perform natural language contracts in uncertain multi-step environments. Within this framework, we develop metrics and baselines for quantifying rational and cooperative play. To evaluate how agents perform at such contracting, we instantiate our framework in ContractSim, an evaluation suite where two players negotiate and execute a multi-turn supplier contract under environmental and inter-player uncertainty. Across six environments and three supplier settings (catering, hotel cleaning, and AI hosting) we find that current LLM-based agents reach agreement reliably, and negotiate efficient contracts when environmental uncertainty is low. However, under high uncertainty, they often fail to negotiate satisfiable, efficient, or mutually beneficial contracts. They are also frequently uncooperative when executing contracts, violating contract terms for additional profit even when contracts are easy to satisfy. These findings highlight room for improvement in the design of language agents that can negotiate, interpret, and execute contracts both rationally and cooperatively.