arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dear Algo:面向统一搜索与推荐的精准优先智能体意图层

Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Hong Yan, Shujian Bu

arXiv 2608.15877首次发表:更新:

发表机构

Meta Platforms Inc.; Google DeepMind(元平台公司; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出面向统一搜索与推荐的精准优先智能体意图层Dear Algo,将多类型意图编译为可执行计划,实验显示其在精度、候选数量及用户相关性上均优于基线。

AI 中文摘要

搜索与推荐服务于共同的发现目标,但对意图的编码方式不同。我们通过Threads平台上的已部署产品Dear Algo研究这一边界,该产品中,“更多NBA新闻”或“减少政治内容”等开放式请求会引导后续信息流推荐,而非返回一次性结果列表。其智能体意图层将显式、推断、负向及复合意图编译为可落地的可执行计划,随后调用传统检索及可选的语义或多模态重排序。该层共享意图到检索的契约,无需在类搜索和类推荐模式间使用单一模型或服务路径。我们在精准优先目标下评估Dear Algo:在对300个公开请求-物品对(296个可评估)的盲审中,严格分类的大语言模型(LLM)评判门控达到94.4%的精确相关精度[88.8%, 98.9%];在72个标准化请求簇中,全配置每20个槽位产生7.73个经评判的候选,而基于LLM生成查询的基线为6.61,提升1.11[0.12, 2.12];在候选随机化服务路径研究(限重排序路径的前72个合格小时)中,经评判准入项的用户加权评判不相关占比为2.80%,而对照为4.78%,差值为-1.97个百分点[-3.02, -0.94],同时精确相关占比高出2.24个百分点[0.08, 4.41]。综上,这些研究展示了在精准优先评估框架下,如何将显式自然语言意图传递至信息流推荐。

英文摘要

Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer subsequent feed recommendations rather than return a one-shot result list. Its agentic intent layer compiles explicit, inferred, negative, and compound intent into a grounded executable plan, then invokes conventional retrieval and optional semantic or multimodal reranking. The layer shares an intent-to-retrieval contract without requiring one model or serving path across search-like and recommendation-like modes. We evaluate Dear Algo under a precision-first objective. In a blinded audit of 300 public request-item pairs (296 evaluable), a strict categorical LLM-as-a-judge gate achieved 94.4\% exact-Relevant precision [88.8\%, 98.9\%]. Across 72 normalized request clusters, the full configuration produced 7.73 judge-qualified candidates per 20 slots versus 6.61 for an LLM-derived-query baseline, a gain of 1.11 [0.12, 2.12]. In a candidate-randomized serving-path study restricted to the reranker path's first 72 eligible hours, the user-weighted judge-Irrelevant share among judged admissions was 2.80\% versus 4.78\% off (-1.97 points [-3.02, -0.94]), while Exact-Relevant share was 2.24 points higher [0.08, 4.41]. Together, these studies show how explicit natural-language intent can be carried into feed recommendation under a precision-first evaluation framework

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑