arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24983cs.CLcs.HCcs.LG

onPanda:通过词元级修正高效标注LLM与智能体的策略内对齐数据

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

  • StepFun(阶跃星辰)
  • Xiamen University(厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

Lei Yang, Mengyin Liu, Jia Wang, Hangyu Guo, Liang Zhao, Zheng Ge, Kang An, Binxing Jiao, Qi Han, Daxin Jiang, Siqi Shen, Xiangyu Zhang

AI总结:

onPanda通过词元级修正交互机制,将LLM对齐数据标注时间中位数降低52%,并生成保留采样分布的策略内数据及细粒度监督信号。

AI中文摘要:

我们提出了onPanda,一个用于高效标注LLM对齐数据和智能体轨迹的交互式工具。onPanda采用词元级修正作为其核心交互方式:在阅读模型响应时,标注者定位第一个不合适的词元,并从模型的候选词元中选择替代词,或通过自由编辑输入正确文本。系统随后截断该位置之后的所有内容,并从修正后的前缀继续生成,重复这种定位-修正-继续的循环,直到获得满意的响应。该机制使标注者能够以低成本精确引导模型输出:一项小型对照研究表明,与手动后编辑相比,onPanda将中位标注时间减少了52%。由于最终响应中的绝大多数词元由模型自身生成,所得数据在很大程度上保留了模型的采样分布,非常适合构建策略内SFT和偏好数据。此外,标注过程中记录的词元级修正提供了具有精确位置的细粒度监督,以及自然配对的正面-负面样本。onPanda还连接到外部工具和框架,支持在真实环境中进行交互式轨迹标注。此外,我们发布了Panda-CVL,一个使用onPanda标注的数据集,以及一个用于词元级修正的基准。

英文摘要:

We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate-correct-continue loop until a satisfactory response is obtained. This mechanism lets annotators precisely steer model outputs at low cost: a small controlled study suggests that onPanda reduces median annotation time by 52% over manual post-editing. Since the vast majority of tokens in the final response are generated by the model itself, the resulting data largely preserves the model's sampling distribution and is well suited for constructing on-policy SFT and preference data. Furthermore, the token-level corrections recorded during annotation provide fine-grained supervision with precise positions and naturally paired positive--negative samples. onPanda also connects to external tools and harnesses, enabling interactive trajectory annotation in realistic environments. In addition, we release Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.

补充信息

↑