arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28605cs.LGcs.AI

UO-FIE:结合精确标签监督与分级效用的实情推断

UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference

Xinchen Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

UO-FIE结合精确标签与分级效用,通过分布预测、软目标和有序损失,在FIE2026微调赛道以0.8316宏效用夺冠。

中文摘要 AI 辅助

Factivity Inference Evaluation 2026(FIE2026)将中文上下文-假设对分类为九个有序的实情区间。其评估指标既奖励精确预测,也奖励接近正确区间的预测,而566个训练样本中有64.1%属于单一类别。在初步实验中,多个mDeBERTa分类模型主要预测主导类别,而Huber回归基线产生的预测更接近正确区间,但精确匹配较少。我们引入了面向效用的实情推断(UO-FIE),一个参数高效的系统,结合了精确标签监督与分级效用。UO-FIE预测九个类别上的分布,并组合硬标签监督、基于效用的软目标、计划类权重和有序损失。我们在受控比较中评估期望效用解码,并使用在袋外预测上选择的有序校准用于提交系统。基于Qwen3.5-9B和LoRA,UO-FIE在微调赛道中以0.8316的宏效用排名第一。一个独立的基于提示的集成在非微调赛道中以0.8450的宏效用排名第三。

英文摘要

The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while 64.1% of the 566 training examples belong to a single class. In preliminary experiments, several mDeBERTa classification models predominantly predict the dominant class, whereas a Huber-regression baseline produces more predictions near the correct interval but fewer exact matches. We introduce Utility-Oriented Factivity Inference (UO-FIE), a parameter-efficient system that combines exact-label supervision with graded utility. UO-FIE predicts a distribution over the nine classes and combines hard-label supervision, utility-based soft targets, scheduled class weights, and an ordinal loss. We evaluate expected-utility decoding in controlled comparisons and use ordinal calibration selected on out-of-fold predictions for the submitted system. Based on Qwen3.5-9B with LoRA, UO-FIE ranks first in the fine-tuning track with a macro utility of 0.8316. A separate prompt-based ensemble ranks third in the non-fine-tuning track with a macro utility of 0.8450.

发表机构

  • Xinjiang University(新疆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑