相同反馈,不同答案:衡量前沿模型客户反馈分析中的运行间不稳定性
Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
浏览论文内容
中文总结 AI 辅助
本研究提出重复运行评估框架,证明分类法接地智能体(TGA)在客户反馈分析中显著降低主题更替和数量分歧,实现更稳定可重复的输出。
中文摘要 AI 辅助
AI智能体正越来越多地被编程用于自动化处理大量非结构化数据的知识工作。这种自动化要求可重复性:当底层证据不变时,智能体的类别、优先级和计数不应在运行之间发生实质性变化,即使每个单独的回答看似合理。我们引入了一个重复运行评估框架,该框架对齐语义等价的类别,并聚焦于两个操作指标:主题更替(返回类别集的归一化变化)和数量分歧(持续存在类别的计数变化)。我们评估了八个前沿模型上的三个重复客户反馈任务,语料库规模从100到5000条记录,多种提示词,以及三种执行设计:原始生成、无分类法的层次分解,以及使用持久主题、子主题和记录级预测的分类法接地智能体(TGA)。在固定使用Claude Opus 4.8和1000条记录语料库的情况下,相对于原始生成和层次分解,TGA将主题更替减少了86%至88%,而匹配主题的数量分歧为零。分类法接地智能体在筛选过程中比每个原始模型都更稳定,在每个语料库规模下都保持更稳定,并且在主题匹配变得更严格或更宽松时保持这一优势。尽管在客户反馈上进行了评估,但该框架更广泛地针对非结构化语料库的重复综合,包括财务报告、法律文件、事件记录和科学文献。总体而言,这些结果表明,分类法接地为重复性知识工作产生了更一致和可重复的输出。
英文摘要
AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, the agent's categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible. We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist. We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions. With Claude Opus 4.8 and the 1,000-record corpus fixed, TGA reduces theme churn by 86--88% relative to both raw generation and hierarchical decomposition, while matched-theme volumes have zero disagreement. The taxonomy-grounded agent is more stable than every raw model in the screen, remains more stable at each corpus size, and keeps this advantage when theme matching is made stricter or looser. Although evaluated on customer feedback, the framework targets repeated synthesis of unstructured corpora more broadly, including financial reports, legal documents, incident records, and scientific literature. Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.