arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26025cs.LGcs.CL

MICRO:用于严重错误发现的多保真度主动搜索

MICRO: Multi-Fidelity Active Search for Severe Error Discovery

Orlando Leone, Niclas Pokel, Pehuén Moure, Yingqiang Gao, Roman Boehringer

首次发表
浏览论文内容

中文总结 AI 辅助

MICRO提出多保真度主动搜索框架,联合建模评分与损失,通过聚类和滚动优化预算分配,在WMT20上显著提升严重错误发现数量。

中文摘要 AI 辅助

人类反馈在成本和信息量上可能有所不同。强反馈能够揭示严重错误,但成本高昂,因此较便宜的质量评分有助于决定哪些项目需要标注。我们提出了MICRO(多保真度影响聚类滚动),一个主动搜索框架,将共享预算分配给这些反馈类型,以最大化确认的严重错误发现数量。MICRO联合建模评分和基于项目特征的标注损失,以指导采集。它根据采集对严重性概率的预测影响进行聚类,以选择多样化的候选,然后使用滚动来估计其发现价值。在WMT20英德上的实验表明,评分改善了损失重建和严重性预测。MICRO在四种预算和评分成本设置中实现了最高的平均发现数量,在一种设置中与改编的MF-ENS表现相似,在其他三种设置中显著优于所有六个比较策略,包括两个滚动对照(p<.001)。

英文摘要

Human feedback can vary in cost and informativeness. Strong feedback can reveal severe errors but is costly, so cheaper quality ratings can help decide which items to annotate. We propose MICRO (Multi-Fidelity Impact Clustered Rollout), an active search framework that allocates a shared budget to these feedback types to maximise confirmed severe error discoveries. MICRO jointly models ratings and annotation losses conditional on item features to steer acquisition. It clusters acquisitions by their predicted impact on severity probabilities to select diverse candidates, then uses rollout to estimate their discovery value. Experiments on WMT20 English-German show that ratings improve both loss reconstruction and severity prediction. MICRO achieves the highest mean discovery count across four budget and rating cost settings, with similar performance to adapted MF-ENS in one and significant gains over all six comparison policies, including two rollout controls, in the other three $(p<.001)$.

发表机构

  • Institute of Neuroinformatics, University of Zurich and ETH Zurich(苏黎世大学和苏黎世联邦理工学院神经信息学研究所)
  • ETH AI Center(苏黎世联邦理工学院人工智能中心)
  • Department of Computational Linguistics, University of Zurich(苏黎世大学计算语言学系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑