arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06815cs.CLcs.AI

类型化联邦工件用于智能体网络:在冻结的异构LLM智能体间共享工具路由知识

Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents

Abhijit Chakraborty, Ni Trieu, Vivek Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

针对开放网络中冻结异构LLM智能体的工具路由知识共享问题,提出类型化联邦工件SYNAPSE1,通过模式验证字段实现隐私与跨模型转移,实验显示其性能接近集中式方案,并显著提升智能体工具调用准确率。

中文摘要 AI 辅助

开放、互联的网络将使智能体能够运行来自多个供应商的冻结模型,保持其历史记录的隐私,并相互教授何时调用哪个工具。扁平文本(提示、示例池)使协议难以区分噪声统计、合并规则和文档。权重和适配器无法在平台之间传递这些知识。我们建议共享类型化联邦工件,即具有明确定义字段的、经模式验证的对象,用于逐字段隐私(此处描述,但已测量)、争议解决和跨模型转移,并将其实例化为SYNAPSE1,一种通用的工具路由知识。在删除192条垃圾条目和1,916条与测试查询重复或几乎重复的训练项后,一个联邦汇编在StableToolBench(3,180个工具)上,每轮每个客户端20 MB JSON的情况下,其路由性能与集中式汇编相差在1.1个点以内。将相同的经验以类型化字段而非单个扁平字符串合并并展示给路由器,在干净数据上价值8.5个点,在60%注入矛盾下价值7.4个点。交叉合并和呈现显示两半是不可分割的(以扁平方式呈现的类型化合并是最差的分支),而三种冲突策略无法区分,因此促使这项工作的冲突日志并非问题所在。在τ-bench零售环境中,每个汇编分支使GPT-4o智能体的每步工具调用准确率至少提高6.7个点,这归因于格式而非联邦经验。论文以两个警示性发现作结:在主题标记的数学代理和StableToolBench上,对相同标记经验进行TF-IDF分类器处理,其性能超过了所有LLM路由分支(分别高出48和26个点,主要在于检索召回率),因为基准的池中包含了每个据称未见工具的标记查询以及我们过滤前每个测试查询的逐字内容。它无法衡量对无标签工具的路径选择,而这正是路由存在的目的。

英文摘要

An open, networked web will allow agents to run frozen models from multiple vendors, keep their history private, and teach each other which tool to call and when. Flat text (prompts, example pools) makes it difficult for the protocol to distinguish between noise statistics, merging rules, and documentation. Weights and adapters cannot transfer that knowledge between platforms. We suggest sharing typed federated artifacts, schema-validated objects with well-defined fields for per-field privacy (described here, but measured), dispute resolution, and cross-model transfer, and instantiating them as SYNAPSE1, a common tool-routing knowledge. After deleting 192 garbage entries and 1,916 training items that duplicate or almost duplicate test queries, a federated compendium routes within 1.1 points of a centralized one at 20 MB of JSON per client each round on StableToolBench (3,180 tools). The same experience merged and shown to the router as typed fields rather than one flat string is worth 8.5 points on clean data and 7.4 under 60% injected contradiction. Crossing merge and rendering shows the halves are inseparable (the typed merge shown flat is the worst arm), while three conflict policies are indistinguishable, so the conflict log that motivated this work is not the On τ-bench retail, each compendium arm improves GPT-4o agents' per-step tool-call accuracy by at least 6.7 points, attributed to format rather than federated experience. Two cautionary findings conclude the paper: on a topic-labeled math proxy and StableToolBench, a TF-IDF classifier over the same labeled experience beats every LLM routing arm (by 48 and 26 points, mostly retrieval recall) because the benchmark's pool holds labeled queries for every supposedly unseen tool and every test query verbatim before our filter. It cannot measure routing to tools without labels, which routing exists for.

发表机构

  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑