arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RE-AD:数据标注的实时需求遵循

RE-AD: Real-Time Requirement Adherence for Data Labeling

Siddarth Malreddy, Ishan Nigam, Akshay Arora, Nikhil Mittal, Subrat Sahu

arXiv 2607.20455首次发表:更新:

发表机构

Uber AI Solutions(优步人工智能解决方案)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对众包数据标注质量问题,引入RE-AD框架,利用大语言模型,通过分解SOP为原子规则、分类并应用分层验证策略来验证标注质量,在合成基准上F1分数达0.749,生产部署中标注者接受并修复大部分错误。

AI 中文摘要

人工标注的数据仍然是训练前沿大语言模型的基础。然而,众包标注往往存在因标注者误解或缺乏参与而导致的质量问题。为解决此问题,我们引入了一个实时需求遵循(RE-AD)框架,该框架利用大语言模型主动验证标注质量。我们的方法包括通过自我反思将标准操作程序(SOP)分解为原子规则,按复杂性对其进行分类,并应用分层验证策略。在一个合成基准上进行评估,该系统的F1分数达到0.749。此外,生产部署使标注者接受并修复了框架标记的82%的错误。我们还进行了消融研究以证明核心设计决策的影响。

英文摘要

Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer from quality issues stemming from annotator misunderstanding or lack of engagement. To address this, we introduce a real-time requirement adherence (RE-AD) framework that leverages LLMs to proactively validate labeling quality. Our methodology involves decomposing Standard Operating Procedures (SOPs) into atomic rules via self-reflection, categorizing them by complexity, and applying tiered validation strategies. Evaluated on a synthetic benchmark, the system achieved an F1 score of 0.749. Furthermore, production deployment resulted in annotators accepting and fixing 82% of the errors flagged by the framework. We include ablation studies to demonstrate the impact of our core design decisions.

CommentsAccepted to The Fifth Generation, Evaluation & Metrics Workshop (GEM) workshop at ACL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑