arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从貌似合理的层级结构到有用的分类体系:评估客户反馈上的智能体工作流

From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback

Prabhath Chellingi, Raviraja G, Viraj Bagal

arXiv 2610.09377首次发表:更新:

发表机构

Enterpret(Enterpret)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出用整树指标评估智能体生成的客户反馈分类体系,发现表面检查通过的树存在高重复命名和跨分支泄漏,强调需同时衡量整树与节点质量。

AI 中文摘要

分类体系是AI系统用来组织证据、聚合模式并在大型文档集合上回答问题的符号表示。在客户反馈领域,类别树决定了每条记录如何被计数和路由、哪些问题会被看到以及由哪个团队负责。智能体工作流现在可以轻松生成一个看起来合理的层级结构,而这类树目前通过通用的、针对单个节点的检查来验证:每个名称符合其描述、位于正确的父节点之下,并且与兄弟节点保持区分。我们提出了一个更具操作性的问题:生成的层级结构何时真正可用作生产级分类体系?我们在两个专有反馈语料库(1,940条和5,000条记录)上构建了六个分类体系:对于每个语料库,有一个生产参考分类体系,以及在同一输入下对同一工作流的两次重复运行。所有六个分类体系都通过了通用的命名和结构检查,更深入的产品覆盖检查甚至更倾向于生成的树。然而,在每个生成的树中,至少有97.7%的叶节点名称仅仅重复了祖先的名称(参考分类体系中分别为13.9%和2.9%),并且在其中一个树中,每四条记录中有三条落在多个顶级类别之下。我们引入了两类整树指标:结构判别器测试树的形状是从数据中学习到的还是由其生成器强加的;团队可划分性测试分支是否将反馈分割成团队可以拥有的组。通用检查评为同样正确的树在跨分支泄漏方面相差27个百分点,并且只有两个树中的一个优于随机分割。表面合理性不足以衡量分类体系质量:评估必须同时衡量整棵树和每个节点。

英文摘要

Taxonomies are the symbolic representations through which AI systems organize evidence, aggregate patterns, and answer questions over large document collections. Over customer feedback, the category tree decides how every record is counted and routed, which problems get seen, and which team owns them. Agentic harnesses now make it easy to generate a plausible-looking hierarchy, and such trees are checked today with generic, individually scoped checks: each name fits its description, sits under the right parent, and stays distinct from its siblings. We ask a more operational question: when is a generated hierarchy actually useful as a production taxonomy? We build six taxonomies over two proprietary feedback corpora (1,940 and 5,000 records): for each corpus, a production reference and two repeated runs of the same harness under identical inputs. All six pass every generic naming and structure check, and a deeper product-coverage check even prefers the generated trees. Yet in every generated tree at least 97.7% of leaf names merely restate an ancestor's name (13.9% and 2.9% in the references), and in one, three of every four records fall under multiple top-level categories. We introduce two families of whole-tree metrics: structural discriminators test whether a tree's shape was learned from the data or imposed by its generator; team partitionability tests whether branches split feedback into groups teams can own. Trees the generic checks rate as equally correct differ by 27 percentage points in cross-branch leakage, and only one of the two beats a random split. Surface plausibility is an insufficient measure of taxonomy quality: evaluation must measure the whole tree as well as each node.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑