arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐私保护的提示策略搜索用于机器人控制

Privacy-Preserving Prompted Policy Search for Robotic Control

Ali Irshayyid, Feng Lin, Chong Li, Jun Chen

arXiv 2609.30554首次发表:更新:

发表机构

Oakland University; Wayne State University; OORT; Columbia University(奥克兰大学; 韦恩州立大学; OORT; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出隐私保护的提示策略搜索(PP-ProPS),通过客户端编码策略参数和奖励,实现LLM引导的策略优化而不泄露专有数据,并在多数控制任务中优于现有方法。

AI 中文摘要

大型语言模型(LLMs)最近在作为强化学习(RL)的上下文策略优化器方面展现出有前景的能力,使得策略搜索能够由数值奖励信号和自然语言推理共同驱动。然而,在实际中部署此类方法需要将原始策略参数和奖励历史传输到基于云的LLM API,从而将专有的控制策略暴露给第三方服务提供商。为解决这一问题,本文引入了隐私保护的提示策略搜索(PP-ProPS),这是一个框架,能够在保持策略和环境参数机密的同时,实现LLM引导的策略优化。PP-ProPS在每次API请求中包含策略参数和奖励值之前,使用秘密的客户端变换对其进行编码,确保LLM提供商仅观察到编码后的策略参数和缩放后的奖励信息。此外,与Vanilla ProPS不同,所提出的框架不需要真实的最优回合回报被LLM知晓或披露。除了保护优化数据外,PP-ProPS还通过两种方式改进搜索过程。首先,它为LLM提供单独的奖励组件,而不仅仅是单一的总回报,为每个候选策略提供更有信息量的反馈。其次,它使用有界的历史记录,防止提示无限增长,从而改善高维策略的搜索,并支持使用开放权重LLM。所提出的PP-ProPS在连续和离散控制问题上进行了评估,涵盖多关节动力学与接触(MuJoCo)运动、经典控制、高速公路驾驶和机械臂操作。与Vanilla ProPS相比,所提出的PP-ProPS在十个评估任务中的七个上优于ProPS,并在六个任务中的五个上超越了包括PPO、SAC和TRPO在内的传统RL方法。

英文摘要

Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy parameters and rewards history to cloud-based LLM APIs, exposing proprietary control strategies to third-party service providers. To address this issue, this paper introduces Privacy-Preserving Prompted Policy Search (PP-ProPS), a framework that enables LLM-guided policy optimization while keeping policy and environmental parameters confidential. PP-ProPS encodes policy parameters and reward values using secret client-side transformations before they are included in each API request, ensuring that the LLM provider observes only encoded policy parameters and scaled reward information. Furthermore, unlike Vanilla ProPS, the proposed framework does not require the true optimal episodic return to be known or disclosed to the LLM. Beyond protecting the optimization data, PP-ProPS improves the search process in two ways. First, it provides the LLM with individual reward components instead of only a single total return, offering more informative feedback about each candidate policy. Second, it uses a bounded history that prevents the prompt from growing indefinitely, improving search with high-dimensional policies and supporting the use of open-weight LLMs. The proposed PP-ProPS is evaluated on both continuous and discrete control problems spanning Multi-Joint dynamics with Contact (MuJoCo) locomotion, classic control, highway driving, and robotic arm manipulation. Compared to Vanilla ProPS, the proposed PP-ProPS outperforms ProPS in seven of the ten evaluated tasks, and surpasses conventional RL methods including PPO, SAC, and TRPO, in five of the six tasks.

Comments9 pages, 5 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑