AI 中文总结
针对现有大语言模型序列标注方法的领域对齐不足与推理效率低问题,提出DIRECT框架,结合训练时优化与推理时校正,在8个数据集上实现性能与效率的显著提升。
AI 中文摘要
序列标注是一种细粒度信息抽取任务,但现有基于大语言模型的方法存在领域对齐不足、推理效率低的问题。为解决这些问题,本文提出DIRECT框架,通过训练时优化与推理时校正处理上述问题。具体而言,DIRECT在监督微调后执行直接偏好优化(DPO)以强化任务与人类偏好的对齐,引入受控解码过程,强制固定输出格式并将预测限制在候选集内;为进一步提升效率,采用模板填充机制,要求模型仅生成标签 token,同时通过键值缓存(KV Cache)复用前缀内容,从而减少冗余计算。在8个数据集上的实验结果表明,与现有方法相比,DIRECT在性能和效率两方面均实现了显著提升。
英文摘要
Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alignment and low inference efficiency. To address these issues, we propose DIRECT, a framework that addresses these issues through training-time optimization and inference-time rectification. Specifically, DIRECT performs Direct Preference Optimization (DPO) after supervised fine-tuning to strengthen task alignment with human preferences, and introduces a controlled decoding process that enforces fixed output formats and restricts predictions to candidate sets. To further improve efficiency, a template-filling mechanism requires the model to generate only label tokens while reusing prefixed content through the KV Cache, thus reducing redundant computation. Experimental results on eight datasets demonstrate that DIRECT achieves significant improvements in both performance and efficiency compared to existing methods.