arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RA-MoWE:用于查询聚类和智能体工作流生成的工作流亲和性嵌入

RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agentic Workflow Generation

Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Haoyu Wang, Haifeng Chen, Xiaowei Jia

arXiv 2610.07851首次发表:更新:

发表机构

Rutgers University; University of Pittsburgh; NEC Labs America; University of Maryland(罗格斯大学; 匹兹堡大学; NEC美国实验室; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RA-MoWE通过工作流亲和性嵌入聚类查询并生成可复用专家工作流,在300个查询测试集上平均得分提升4.04个百分点,推理调用减少27.7%。

AI 中文摘要

智能体工作流使大型语言模型(LLMs)能够通过协调推理、工具使用和验证来解决复杂任务。然而,针对整个任务集合优化的工作流可能会忽略单个查询所需的推理策略差异,而为每个查询搜索新工作流则会重复昂贵的优化过程。为解决这一权衡问题,我们引入了RA-MoWE,一个利用工作流亲和性嵌入对查询进行聚类并指导可复用专家工作流生成的框架。每个嵌入记录了固定参考工作流集合解决某个查询的效果,揭示了哪些推理策略有效的相似性。RA-MoWE利用每个聚类的查询和平均嵌入,通过执行反馈来初始化和细化专门的工作流。一个嵌入编码器从查询文本预测这些嵌入,使新查询无需先执行参考工作流即可选择生成的专家。在从涵盖数学、科学和编程的四个基准中抽取的300个查询测试集上,RA-MoWE在参考工作流中选择的基础上,将平均任务得分提高了4.04个百分点,同时在推理时减少了27.7%的语言模型调用。

英文摘要

Agentic workflows enable large language models (LLMs) to solve complex tasks by coordinating reasoning, tool use, and verification. However, a workflow optimized for an entire task collection can overlook differences in the reasoning strategies that individual queries need, while searching for a new workflow for every query repeats costly optimization. To address this tradeoff, we introduce RA-MoWE, a framework that uses workflow-affinity embeddings to cluster queries and guide the generation of reusable expert workflows. Each embedding records how well a fixed set of reference workflows solves a query, revealing similarities in which reasoning strategies are effective. RA-MoWE uses each cluster's queries and average embedding to initialize and refine a specialized workflow through execution feedback. An embedding encoder predicts these embeddings from query text, allowing new queries to select a generated expert without first executing the reference workflows. On a 300-query test set drawn from four benchmarks spanning mathematics, science, and programming, RA-MoWE improves average task score by 4.04 percentage points over selecting among the reference workflows, while using 27.7% fewer language-model calls at inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑