发表机构
MetaAI(MetaAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体在社交决策中忽视潜在关系的问题,提出ReAdapt框架,通过显式结构化社交状态和策略驱动的适应步骤,在合成社交世界基准上显著提升反应选择与热忱引荐的准确率。
AI 中文摘要
社交智能体的最基本决策(我是否应该回复这条帖子?我应该联系谁?)并非纯粹的内容问题。正确的行动往往取决于人与人之间潜在的关系——关系强度、互惠性、共同联系——而非哪种内容最为突出。标准的LLM智能体循环并未明确表示新的关系证据应如何修正智能体当前的社会假设,这使得它们在关系线索与内容线索相背离时容易做出表面明显的选择。我们通过一个关系推理基准来形式化这种失败模式:500个包含友谊、关注、反应历史和信息流的合成社交世界,产生1000个查询,涵盖两个任务:反应选择和热忱引荐(寻找到达目标人物的最佳桥梁)。按构造,表面明显的候选与基于关系的基准答案在大约53%的查询中不同,形成了一个颠覆子集,其中智能体必须使用关系证据来修正最初看似合理的选择。我们提出了ReAdapt(关系自适应智能体与策略驱动状态),它通过一个显式的结构化社交状态z = (G, B, R, N, D)来增强ReAct循环,该状态捕获目标、信念、关系、规范和披露。在每次工具观察后,ReAdapt运行一个类型化的适应步骤,更新此状态并发出策略操作(继续、切换、放弃或澄清),然后再选择下一个动作。使用Gemini-3-Flash在每任务n = 150个查询的分层子集上,ReAdapt将热忱引荐准确率从37%提高到51%(+14个百分点),反应选择准确率从69%提高到77%(+8个百分点)。Oracle遗憾分别从0.260降至0.152,从0.095降至0.053。在保持模型、工具和环境不变的情况下,这些结果表明,显式的关系状态适应有助于LLM智能体将检索到的社会证据转化为修正后的决策。
英文摘要
A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient. Standard LLM agent loops do not explicitly represent how new relational evidence should revise the agent's current social hypothesis, leaving them prone to surface-obvious choices when relational and content cues diverge. We formalize this failure mode with a relationship-reasoning benchmark: 500 synthetic social worlds with friendships, follows, reaction histories, and feeds, yielding 1,000 queries over two tasks, reaction selection and warm introduction (finding the best bridge to a target person). By construction, the surface-obvious candidate differs from the relationship-grounded oracle in about 53% of queries, forming an overturn subset where the agent must use relational evidence to revise an initially plausible choice. We propose ReAdapt (Relationship-Adaptive Agent with Policy-driven sTate), which augments the ReAct loop with an explicit structured social state z = (G, B, R, N, D) capturing goal, belief, relationship, norm, and disclosure. After each tool observation, ReAdapt runs a typed Adapt step that updates this state and emits a policy operation (continue, switch, abandon, or clarify) before choosing the next action. With Gemini-3-Flash on a stratified subset of n = 150 queries per task, ReAdapt improves warm-introduction accuracy from 37% to 51% (+14 points) and reaction-selection accuracy from 69% to 77% (+8 points). Oracle regret drops from 0.260 to 0.152 and from 0.095 to 0.053, respectively. Holding the model, tools, and environments fixed, these results suggest that explicit relational-state adaptation helps LLM agents turn retrieved social evidence into revised decisions.
Comments12 pages