arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Goldsmith:基于金损失引导的定义优化与智能体标注框架

Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness

Yihan Li, Hanyi Zhang, Xiaoxi Jiang, Man Guo

arXiv 2610.09489首次发表:更新:

发表机构

Sun Yat-sen University; South China University of Technology(中山大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Goldsmith利用小型金集通过可执行损失引导的智能体优化生成结构化标注定义,在多个任务上优于现有提示优化方法,实现稀缺专家监督下的可扩展标注。

AI 中文摘要

许多标注项目在专家尚未拥有稳定的指南或足够的标签来训练特定任务模型之前就已启动。我们提出了Goldsmith,一个智能体流水线,将一个小型金集——代表预期任务边界的专家标注校准示例——转化为可复用的结构化标注定义。Goldsmith将此定义视为一个可训练的文本来对象。候选定义在相同的金示例上运行,并通过可执行的结构化损失进行评分,而输出模式、格式化、检索、修复、评判和人工审查则保留在外部框架中。大型语言模型(LLM)编辑器将损失最高的失败转化为文本梯度修订,仅在测量损失降低时才接受这些修订。在提示优化比较中,Goldsmith在匹配的评估协议下优于直接重写、OPRO、APE和PromptBreeder。当与检索、基于分数的路由和人工审查相结合时,所得定义还能改善跨类型跨度、成对关系及固定触发事件-参数任务的下游标注。这些结果表明,稀缺的专家监督可以同时支持任务定义学习和可扩展的标注。

英文摘要

Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.

Comments20 pages, 4 figures, 11 tables. Accepted to the main conference of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑