发表机构
Mila – Quebec AI Institute; Skyfall AI; McGill University(米拉-魁北克人工智能研究所; 天空fall人工智能公司; 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出语义老虎机框架,发现LLM的探索-利用受语义先验偏置,语义标签对齐奖励可提升性能,负奖励比同等正奖励触发更多探索,为LLM智能体决策可靠性提供了新见解。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署为需要复杂环境探索的场景中的决策智能体。然而,现有研究对LLMs如何实际平衡探索与利用提出了疑问。与经典智能体不同,LLM智能体通过自然语言与任务交互,使它们接触到任务结构中无正式对应物的语义信息。我们引入语义老虎机(semantic bandit),这是多臂老虎机设置的扩展,明确考虑分配给动作的文本标签,并使用它来研究语义先验——预训练期间从语言与预期奖励的关联中习得的归纳偏置——如何塑造LLM的探索行为。我们发现,语义信息丰富的动作标签会减少探索,转而偏向利用;当与奖励结构对齐时,这会提升性能,而当对齐错误时,会严重降低性能。我们进一步发现,与同等的正奖励相比,负奖励会触发多得多的探索,这与预训练数据中常见的奖励约定所诱导的预期规模偏置一致。总体而言,我们认为,使用语言来定义环境和奖励会引入不可避免的偏置,因为模型是基于词共现进行训练的,这对LLM智能体在现实世界决策场景中的可靠性和稳健性具有影响。
英文摘要
Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic priors --- inductive biases arising from associations between language and expected reward learned during pre-training, shape LLM exploration behaviour. We find that semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned. We further find that negative rewards trigger substantially more exploration than equivalent positive rewards, consistent with an expected-scale bias induced by reward conventions common in pre-training data. Overall, we argue that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.
Comments10 pages, 5 figures in main body