发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究为CheckThat! 2026任务3开发UTS系统,采用HostCite和ShadowVal干预措施,以Llama-3.2:1B为引用验证器,在11支队伍中获第2名,M4得分0.484,明确了引用相关优化方向。
AI 中文摘要
CheckThat! 2026的任务3要求系统生成事实核查文章,评估指标为四个子指标的未加权平均值(M4)。我们的UTS提交作品在11支队伍中排名第2,M4得分为0.484。所部署的系统是一个确定性的草稿生成器,辅以两个单杠杆干预措施:领域归属引用框架(HostCite)和经影子验证的锚点选择器(ShadowVal),仅将Llama-3.2:1B用作每个引用的验证器,而非正文生成器。该系统在WatClaimCheck验证集上比基础草稿将M4提升了0.027,在蕴含度和覆盖度上优于其他参赛队伍,且遵循了我们的消融矩阵明确得出的两条设计规则。评分者保守性:仅对参考文献所蕴含的内容给予分数,模板不计分;大语言模型生成的正文、审稿人姓名及原始证据均不计分。辅助锚点信号与Llama评判器存在校准偏差:我们尝试的所有锚点代理(交叉编码器、长度、开头位置)所选锚点均被评判器拒绝,需以评判器本身作为闸门。与冠军队伍相差的剩余0.062分差距在于引用的精确率和召回率(0.299 vs 0.671),这与放弃低置信度引用的选择性输出策略一致。
英文摘要
CheckThat! 2026 Task 3 asks systems to generate fact-checking articles, graded by an unweighted mean of four sub-metrics (M4). Our UTS submission placed 2nd of 11 teams (M4 = 0.484). The shipped system is a deterministic stub drafter wrapped by two single-lever interventions: a domain-attribution cite frame (HostCite) and a shadow-validated anchor picker (ShadowVal) that use Llama-3.2:1B only as a per-cite validator, never as a body-prose generator. The stack lifts M4 by +0.027 over the stub on the WatClaimCheck validation split, beats the field on entailment and coverage, and follows two design rules our ablation matrix made unambiguous. Scorer conservatism: credit only tokens the references entail - templates pay; LLM prose, reviewer names, and raw evidence all fail. Auxiliary anchor signals are miscalibrated against the Llama judge: every anchor proxy we tried (cross-encoder, length, lead position) picks anchors the judge rejects - gate on the judge itself. The remaining +0.062 gap to the winner sits on citation precision/recall (0.299 vs 0.671), consistent with a selective-emission policy that drops low-confidence cites.
CommentsCLEF2026, CheckThat!2026