arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单独对齐,共同错位:预测大语言模型智能体群体中的对抗捕获

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

Isotta Magistrali, Chen Shani

arXiv 2608.22444首次发表:更新:

AI 中文总结

该研究针对LLM智能体群体,提出从群体无攻击时的良性运行状态校准响应函数,可提前预测坚定少数群体引发的对抗捕获,且捕获为临时状态,孤立对齐不等同于群体对齐。

AI 中文摘要

当前AI安全评估的单位仍为单个模型,但语言模型智能体越来越多地以相互读取和执行彼此决策的群体形式部署,这提出了单个智能体审计无法回答的问题:一个自身校准良好的智能体,仍可能被周围的智能体拉向不同的决策方向。我们在安全分诊任务中研究这一问题,在此任务中,语言模型监控器群体决定是否升级或驳回警报,我们可向其中注入一个始终朝某一方向推动的坚定少数群体。我们发现,单个智能体自身判断几乎相同的两个警报,可使群体行为产生巨大差异,因此审计任何单个成员都无法揭示群体的实际行为。然而这种群体行为可提前预测:仅从群体无攻击者的良性运行状态,我们就能校准一个响应函数,在任何攻击实施前预测坚定少数群体将使群体偏离的程度。随后我们探究改变结果的因素,发现让智能体查看彼此的推理可中和较弱的攻击,但仅能延迟较强的攻击,使问题从群体是否会收敛到攻击者的选择转变为何时收敛。最后我们排除了捕获是不可逆陷阱的假设:一旦移除坚定智能体,群体将漂移回初始状态,因此捕获是一种临时状态。孤立状态下的对齐不等同于群体中的对齐,但群体在攻击下的行为可在任何攻击者出现前,从其之前的行为中提前读取。

英文摘要

The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent audit can answer: an agent that is well-calibrated on its own may still be pulled toward a different decision by the agents around it. We study this on a security-triage task, where populations of language-model monitors decide whether to escalate or dismiss alerts, and into which we can inject a committed minority that always pushes one way. We find that two alerts a single agent judges almost identically on its own can drive collective behavior far apart, so auditing any one member need not reveal what the population will do. Yet that collective behavior can be predicted in advance. From a population's benign, adversary-free operation alone, we calibrate a response function that forecasts, before any attack is run, how far a committed minority will later move it. We then ask what shifts the outcome and find that letting agents see each other's reasoning neutralizes a weak attack, while only delaying it against a strong one, turning the question from whether the population converges on the adversaries' choice into when. Finally, we exclude the hypothesis of capture being an irreversible trap: once the committed agents are removed, the population drifts back toward where it began, so capture is a temporary state. Alignment in isolation is not alignment in a population, yet what a population will do under attack can be read in advance, from how it behaves before any adversary arrives.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑