arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少说谎,说大谎:针对LLM机器人团队的隐蔽内部攻击

Lie Rarely, Lie Big: Stealthy Insider Attacks on LLM Robot Teams

Sribalaji C. Anand, George J. Pappas

arXiv 2610.04744首次发表:更新:

发表机构

University of Pennsylvania; KTH Royal Institute of Technology(宾夕法尼亚大学; 瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究LLM机器人团队中单个被攻陷机器人通过隐蔽谎言破坏共同成果的威胁,推导出验证概率下界和地图误差上界,并实验验证了界限的有效性。

AI 中文摘要

当机器人团队将规划和相互信任委托给LLM代理时,一个被攻陷的机器人就能破坏共同的成果。我们在一个具体任务中研究这一威胁:一项多机器人调查,其中测量结果可以对照物理世界进行验证,但每次验证都会消耗本可用于推进任务的预算。我们将被攻陷的机器人视为系统理论意义上的隐蔽对手:它不受能量界限的限制,而是受团队自身检测器的限制。然后我们推导出两个界限。首先,对手报告被验证的概率下界由通信图中的度数和验证预算决定。其次,任何隐蔽对手造成的地图误差上界由对手偏差分布上的线性规划值决定;其解是隐蔽预算与损害之间的兑换率:在临界验证水平以下,最严重的隐蔽攻击会在最不可能被验证的记录上发出罕见的、全幅度的谎言;在临界验证水平以上,更好的选择是隐藏在噪声中的小偏差。在诚实机器人为LLM代理的实验中,两个界限在攻击实际花费的预算下均成立。实验还表明,LLM机器人重新检查哪些记录是无偏的,但重新检查的程度是不可预测的。

英文摘要

When a team of robots delegates planning and mutual trust to LLM agents, a single compromised robot can corrupt the shared outcome. We study this threat in a grounded task: a multi-robot survey in which measurements can be verified against the physical world, but every verification costs budget that would otherwise advance the mission. We treat the compromised robot as a stealthy adversary in the system-theoretic sense: it is limited not by an energy bound but by the team's own detectors. We then derive two bounds. First, the probability that the adversary's reports are verified is bounded below in terms of the degrees in the communication graph and the verification budget. Second, the map error caused by any stealthy adversary is bounded above by the value of a linear program over the adversary's bias distributions; its solution is an exchange rate between stealth budget and damage: below a critical verification level the worst stealthy attack tells rare, full-magnitude lies on the records least likely to be verified, and above it the better purchase is small biases hidden in the noise. In experiments where the honest robots are LLM agents, both bounds hold at the budget the attack actually spent. The experiments also show that which records an LLM robot re-checks is unbiased, but how much it re-checks is unpredictable.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑