arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01597cs.CLcs.AI

语言强化学习的兴起

The Rise of Verbal Reinforcement Learning

Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu

首次发表
浏览论文内容

中文总结 AI 辅助

该文提出语言强化学习(VRL)范式,围绕语言反馈的生效时机与修改对象划分三大支柱,阐述其对智能体开发的重塑作用及相关挑战机遇。

中文摘要 AI 辅助

自然语言正成为改进语言智能体的主要反馈渠道,能够以人类和现代语言模型均可理解的形式传达意图、偏好和因果结构。我们将这一范式称为语言强化学习(Verbal Reinforcement Learning,VRL),并提供对它的首个统一阐述。我们围绕单一轴——语言反馈在智能体生命周期中生效的时机及其修改对象——来组织该领域,形成三大支柱:(1)语言作为基础信号,即语言通过指定目标、状态和奖励结构来定义任务本身;(2)语言作为审议反馈,即自然语言在测试时引导推理,无需更新模型参数;(3)语言作为学习信号,即基于语言的反馈通过训练来调整模型参数。在每个支柱内,我们综合代表性研究,区分方法的关键子类别,并概述语言在塑造智能体行为中所起的不同作用。这一分类法表明,语言强化正如何重塑智能体开发,同时也为构建更强大、更对齐的智能体定义了挑战与机遇。

英文摘要

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and reward structures; (2) \textbf{Language as Deliberative Feedback}, where natural language guides reasoning at test time without the need to update model parameters; (3) \textbf{Language as Learning Signal}, where language-based feedback shapes model parameters through training. Within each pillar, we synthesize representative work, distinguish key subcategories of approaches, and outline the distinct role language plays in shaping agent behavior. Together, this taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.

发表机构

  • Capital One
  • University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

↑