arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38093cs.AI

风险厌恶型智能体的性格训练

Character Training for Risk-Averse Agents

  • Columbia University(哥伦比亚大学)
  • UK AI Security Institute(英国人工智能安全研究所)
  • Arcadia Impact
  • National University of Singapore(新加坡国立大学)
  • Resolution

机构由 AI 辅助整理,请以论文原文为准。

Arav Dhoot, Punya Syon Pandey, Jamie Johnson, Daniel Tan, Elliott Thornley, David Demitri Africa

AI总结:

通过性格训练向智能体灌输风险厌恶偏好,使其倾向安全策略,实验表明该方法有效且可扩展,有助于缓解错位AI风险。

AI中文摘要:

资源方面的风险厌恶可以防止错位的AI智能体造成灾难性伤害。错位但风险厌恶的智能体倾向于选择更安全的策略,例如与人类达成协议,而不是更冒险的策略,如反抗。我们通过性格训练使智能体变得风险厌恶,发现人格特质为灌输风险偏好提供了稳健的机制。为此,我们构建了一个模型章程,描述智能体资源上的常数绝对风险厌恶(CARA),并通过在策略蒸馏来灌输这一特性。尽管在训练期间从未见过基准测试的决策格式,经过性格训练的模型与直接在该格式上训练的基线模型表现相当,并且在我们的四个模型中,有两个模型在分布外泛化方面优于基线。我们还调整了章程的不同方面,发现令牌预算和模型选择是性格训练中灌输风险厌恶最有影响力的方面。我们从这些结果得出结论,性格训练是一种有前景且可扩展的方式,用于灌输广泛的倾向,我们可以利用这一点来减轻错位AI智能体带来的风险。

英文摘要:

Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.

↑