发表机构
Kaons 𝑲 ∗(Kaons 𝑲 ∗)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对政治逃避检测难题,提出AsymVerify系统用于SemEval-2026任务6的三元分类。核心方法是对问答对分类后对低置信预测进行非对称验证,主要贡献是取得高排名,在不同数据集上提升宏F1分数,且降低推理成本。
AI 中文摘要
政治逃避难以检测,因为逃避性回答往往看似合作却避免具体承诺。我们提出了AsymVerify,这是一个用于SemEval-2026任务6的置信门控验证系统,该任务对明确回复、矛盾和明确无回复进行三元分类。AsymVerify在评估集(D_eval,n = 237)上的宏F1分数为0.85,在官方排行榜的41支队伍中排名第二。该系统先对每个问答对进行分类,然后对低置信度预测选择性地应用降级验证(CR/CNR -> AMB)或升级验证(AMB -> CR)。开发分析表明错误集中在矛盾边界的两个方向,这促使了这种非对称双验证器设计,同时置信门控使额外推理成本较低。在D_dev(n = 308)上,使用GLM-4.7的AsymVerify在每次示例1.48次调用时比单遍分类获得了+17.1的宏F1,仅升级验证器就在D_dev上使每个测试的LLM后端比其单遍基线提高了+6.8到+15.2的宏F1。代码可在这个https URL获取。
英文摘要
Political evasion is difficult to detect because evasive answers often appear cooperative while avoiding concrete commitment. We present AsymVerify, a confidence-gated verification system for SemEval-2026 Task 6, a three-way classification of Clear Reply, Ambivalent, and Clear Non-Reply responses. AsymVerify scored 0.85 Macro F1 on the evaluation split (D_eval, n=237), placing 2nd out of 41 teams on the official leaderboard. The system first classifies each question-answer pair, then selectively applies downgrade verification (CR/CNR -> AMB) or upgrade verification (AMB -> CR) to low-confidence predictions. Development analysis shows that errors concentrate at the Ambivalent boundary in both directions, motivating this asymmetric two-verifier design while confidence gating keeps additional inference cost low. On D_dev (n=308), AsymVerify with GLM-4.7 gains +17.1 Macro F1 over single-pass classification at 1.48 calls/example, and the upgrade verifier alone improves every tested LLM backend on D_dev by +6.8 to +15.2 Macro F1 over its single-pass baseline. Code is available at https://github.com/kaons-research/AsymVerify-ACL.
CommentsAccepted to SemEval-2026 Task 6 (CLARITY) at ACL 2026. Team AsymVerify placed 2nd of 41 teams on the official Subtask 1 leaderboard. Task: https://konstantinosftw.github.io/CLARITY-SemEval-2026/. Code: https://github.com/kaons-research/AsymVerify-ACL. Website: https://kaons.com/