评估即搜索:会议助手接地故障的自适应发现
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出EaS方法构建MeetingProbe基准,发现会议助手故障多源于话语语用挑战,其自适应搜索比随机探测的故障发现率高2.5倍,且公开了该基准以支持可复现评估。
AI中文摘要:
由大语言模型(LLM)驱动的会议助手已大规模部署,但对其接地保真度的系统评估仍局限于静态基准,无法覆盖与特定话语结构或推理需求相关的故障模式。本文提出评估即搜索(Evaluation-as-Search, EaS),一种反馈驱动的方法,将质量评估框架化为对会议参与者可能提出的自然问题空间的自适应搜索。EaS不采用均匀采样,而是在迭代过程中从评估者反馈中学习,结合基于UCB评分的覆盖图与盲多维质量评估,将探测精力集中在最可能出现故障的认知需求上。利用EaS,我们构建了MeetingProbe基准,包含3000多个带注释的问答对,覆盖三种会议类型的20份 transcript 及三种LLM助手。消融实验显示,自适应搜索的故障发现率是随机探测的2.5倍(7.1% vs 2.9%),其中策略规划器贡献最大的个体效应。在三个模型上,我们观察到清晰的能力梯度,并识别出八大反复出现的故障类别,以话语语用挑战为主,而非事实回忆错误。我们进一步在多个模型家族和提供商上验证了MeetingProbe,发现了清晰的能力梯度及一组所有模型均无法处理的通用故障精选子集。MeetingProbe已公开发布,以支持对会议助手接地保真度的可复现评估。
英文摘要:
LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation-as-Search (EaS), a feedback-driven methodology that frames quality evaluation as an adaptive search over the space of natural questions a meeting participant might ask. Rather than sampling uniformly, EaS learns from evaluator feedback across iterations to concentrate probing effort on cognitive demands where failures are most likely, guided by a UCB-scored coverage map and blind multi-dimensional quality evaluation. Using EaS, we construct MeetingProbe, a benchmark of over $3{,}000$ annotated question--answer pairs spanning 20 transcripts from three meeting genres and three LLM assistants. In ablations, adaptive search surfaces $2.5\times$ more failures than random probing ($7.1\%$ vs. $2.9\%$ finding rate), with the strategic planner contributing the largest individual effect. Across three models, we observe a clear capability gradient and identify eight recurring failure categories dominated by discourse-pragmatic challenges rather than factual recall errors. We further validate MeetingProbe across multiple model families and providers, finding a clean capability gradient and a curated subset of universal failures that no model handles. MeetingProbe is released publicly to support reproducible evaluation of meeting assistant grounding fidelity.