发表机构
Tencent Inc; Beijing Institute of Technology; University of Science and Technology of China(腾讯公司; 北京理工大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有智能搜索框架在复杂社交搜索中性能下降的问题,提出适配社交场景的智能搜索框架SocialBuddy,构建大规模模拟环境SocialEnv与混合粒度优化框架SocialPO,经实验其35B参数版本性能远超更大规模前沿大语言模型。
AI 中文摘要
在数字社交互动时代,从海量社交信息流中搜索好友的帖子已成为用户的基本需求。然而,尽管现代智能搜索框架在常规检索任务中取得了显著成功,但面对异构用户查询和多维社交信息流时,它们会失效,导致复杂社交搜索任务的性能严重下降。为弥合这一差距,我们推出了SocialBuddy,这是首个适配社交场景的智能搜索框架。具体而言,我们构建了SocialEnv,这是首个用于社交搜索的大规模模拟环境。借助自动化的数据与轨迹合成流程,SocialEnv包含20万个用户画像、1000万条社交帖子和5万条推理轨迹,为社交搜索智能体的开发奠定了坚实基础。为解决社交搜索中稀疏奖励导致的信用分配难题,我们设计了SocialPO,这是一个混合粒度优化框架。它通过多维奖励在宏观层面强化成功的推理路径,同时通过细粒度前缀截断和token级监督在微观层面修正偏离的轨迹。这种混合粒度设计为复杂长序列场景提供了多尺度指导。最后,我们构建了SocialSearch基准,为评估SocialBuddy的社交搜索能力提供量化评估方案。大量实验表明,SocialBuddy-35B的性能远超规模大得多的前沿大语言模型(LLM)。代码和数据集将在论文接收后发布。
英文摘要
In the era of digital social interaction, searching friends' posts from massive social streams has become a fundamental user need. However, while modern agentic search frameworks have achieved remarkable success in conventional retrieval tasks, they break down when confronted with heterogeneous user queries and multi-dimensional social feeds, resulting in severe performance degradation in complex social search. To bridge this gap, we introduce SocialBuddy, the first agentic search framework tailored for social scenarios. Specifically, we construct SocialEnv, the first large-scale simulated environment for social search. Powered by an automated data and trajectory synthesis pipeline, SocialEnv includes 200K user profiles, 10 million social posts, and 50K reasoning trajectories, establishing a solid foundation for the development of social search agents. To tackle the credit assignment dilemma caused by sparse rewards in social search, we design SocialPO, a hybrid-granularity optimization framework. It macroscopically reinforces successful reasoning paths via multi-dimensional rewards, while microscopically rectifying deviated trajectories through fine-grained prefix truncation and token-level supervision. This hybrid-granularity design delivers multi-scale guidance in complex long-sequence scenarios. Finally, we construct SocialSearch Benchmark to provide a quantitative evaluation scheme for assessing the social search capabilities of SocialBuddy. Extensive experiments demonstrate that SocialBuddy-35B surpasses significantly larger frontier LLMs. Code and dataset will be released upon article acceptance.