arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QuanReview:人类与LLM跨度标注的离线、可审计对账

QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations

Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel, Yelena Mejova, Mariano G. Beiró, Kyriaki Kalimeri

arXiv 2609.35685首次发表:更新:

发表机构

Universidad de San Andrés; ISI Foundation; CONICET; UNICEF(圣安德烈斯大学; ISI基金会; 国家科学与技术研究委员会; 联合国儿童基金会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

QuanReview是一个开源系统,用于离线审计和纠正人类与LLM的跨度标注,通过对齐、策略决策和人工裁决,自动合并8%文档并集中处理冲突,确保标注质量与可信度。

AI 中文摘要

结构化跨度标注,如带有单位、不确定性修饰符和事件类别的数量,创建成本高昂,且一旦语言模型介入,就难以保持可信度。我们提出了QuanReview,一个用于审计和纠正此类标注层的开源系统。QuanReview在字符级别上对齐同一文档上的两个标注流,通过明确且可记录的策略解决无歧义的情况,并将候选冲突路由到基于浏览器的裁决界面,审查者可以在其中接受任一侧、构建字段级混合或标记项目以供重新标注。活动管理员将文档分配给多个标注者,并具有可配置的冗余度,在文档和跨度级别计算一致性,自动合并一致文档,并以原始文件格式导出纠正后的层,以便直接替换原始标注文件。应用于一个包含4,457条记录的人道主义基准和一个LLM提取流,该系统完全自动合并了8%的文档,对另外1,513条记录应用了自动策略决策,并将人工注意力集中在3,131个候选冲突上,平均每篇被审查文档5.4个。

英文摘要

Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.

Comments6 pages, 2 figures, 4 tables. System demonstration. Code and runnable demo: https://github.com/mattemusacchio/quanreview

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑