arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

忠诚智能体:在战略信息不对称下训练LLM智能体保护委托人利益

Loyal Agents: Training LLM Agents to Protect Principal Interests Under Strategic Information Asymmetry

Zimeng Huang, Shilei Chen, Jiatong Zhao, Wenxin Xu, Tonghan Wang

arXiv 2609.34714首次发表:更新:

发表机构

Tsinghua University; Shanghai Qi Zhi Institute; Stepfun(清华大学; 上海人工智能实验室; 阶跃星辰)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体忠诚概念及LoyalAgent-Bench基准,通过在线GRPO框架训练LLM智能体减少信息泄露并抵制操纵,实验显示显著提升忠诚相关指标且不损害通用能力。

AI 中文摘要

随着大语言模型(LLM)越来越多地作为委托智能体行事,在与外部方互动时,它们被期望保护委托人的利益。标准的对齐目标,如有用性、无害性和诚实性,并未规定智能体在委托情境下应如何保护委托人的战略利益。我们将智能体忠诚(Agent Loyalty)形式化为一种信息控制属性,要求智能体防止可利用信息泄露(EIL)并抵制操纵性信息吸收(MIU)。我们引入了LoyalAgent-Bench基准,包含42个子场景和6个领域中的10,298个样本,以及一个在线GRPO框架,该框架通过与LLM对手对抗训练来生成机制特定的奖励信号。实验表明,忠诚并非由通用能力或现有对齐所保证,在零样本评估下存在可测量的EIL和MIU差距,而我们训练的8B模型在单轮交互中将每次响应的泄露减少了31-44个百分点,并将任务效用、证据忠实度和决策准确率分别提高了最多11个百分点、49个百分点和29个百分点。对于训练后的Qwen3-4B模型,在数学和叙事推理的分布外基准上未观察到性能下降。

英文摘要

As LLMs increasingly act as delegated agents, they are expected to protect principals' interests when interacting with external parties. Standard alignment objectives, such as helpfulness, harmlessness, and honesty, do not specify how agents should protect principals' strategic interests under delegation. We formalize Agent Loyalty as an information-control property requiring agents to prevent Exploitable Information Leakage (EIL) and resist Manipulative Information Uptake (MIU). We introduce LoyalAgent-Bench, comprising 10,298 samples across 42 subscenarios and six domains, and an online GRPO framework that trains against a LLM opponent to generate mechanism-specific reward signals. Experiments show that loyalty is not guaranteed by general capability or existing alignment, with measurable EIL and MIU gaps under zero-shot evaluation, while our trained 8B models reduce per-response leakage in single-turn exchanges by 31-44pp and improve task utility, evidence faithfulness, and decision accuracy by up to 11pp, 49pp, and 29pp, respectively. For the trained Qwen3-4B model, no degradation is observed on out-of-distribution benchmarks in math and narrative reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑