arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

异构偏好下基于显式人格推断的大语言模型编排

Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

Shuqing Shi, Ziyan Wang, Milind Tambe, Yali Du

arXiv 2610.07587首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HARP框架,通过显式贝叶斯后验推断智能体私有偏好,解决LLM编排中信念更新缺失问题,实现最优遗憾并增强探索。

AI 中文摘要

大语言模型编排研究的是编排器如何协调一组自主智能体以实现共同目标或最大化集体福利。这些智能体通常是异构的,每个智能体都持有私有的偏好,且不会公开。从行为中推断这类隐藏偏好一直是博弈论和多智能体系统中长期研究的课题。核心挑战在于对每个智能体的偏好维持一个信念,并根据观察到的智能体行为更新该信念。现有的大语言模型编排器将该信念作为提示文本携带,没有显式的更新规则,这导致早期错误持续存在并传播而非被纠正。为此,我们提出HARP(异构偏好智能体通过偏好推断进行编排),一种新颖的框架,将信念从提示中移出。具体而言,HARP为每个智能体在有限候选偏好集上维护一个数值后验,并通过贝叶斯规则以闭式更新。语言模型仅提供动作和每个候选的似然,因此估计与其推理解耦。我们证明,当因子分解精确时,HARP达到与显式联合推断相同的O~(√K)贝叶斯遗憾。此外,HARP+在规划中增加对区分候选的动作的奖励,使得即使在最优动作无信息时推断也能继续。在三个基底上的实证结果,从偏好完全决定收益,到收益依赖于更多因素,再到显式联合推断不可行的规模,表明HARP+在我们理论识别的类别中是最强的非预言方法。

英文摘要

LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same $\tilde O(\sqrt K)$ Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.

CommentsThere are confusions on the preferences and personas in the introductions and also misunderstanding in the title

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑