LLM-INSTRUCT在UZH 2026共享任务中的应用:用于段落级论证挖掘的约束感知检索和选择性辩论
LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining
浏览论文内容
中文总结 AI 辅助
本文介绍了在UZH 2026共享任务中用于段落级论证挖掘的获胜系统LLM-INSTRUCT,通过元数据感知检索、约束解码等方法缩小决策空间,提高了准确性和提交稳健性,在官方排行榜上取得优异成绩。
中文摘要 AI 辅助
我们展示了LLM-INSTRUCT,这是在2026年ArgMining关于联合国和教科文组织决议中段落级论证挖掘的UZH共享任务中的获胜系统。该任务需要进行段落类型分类、预测141个官方标签的子集以及在严格的JSON模式设置下仅使用参数高达8B的开放权重模型进行定向关系预测。我们将该任务构建为约束结构化预测。系统首先通过元数据感知密集检索缩小候选标签空间,然后应用带维度上限的约束解码,仅将不确定情况升级到三主体辩论分支,最后验证输出模式。在官方排行榜上,LLM-INSTRUCT总体排名第一,F1得分第一,在LLM作为评判方面排名第五。在开发过程中,我们的配置搜索将任务1b的微F1从35.83%进一步提高到40.08%,同时保持内部任务2得分在4.421。主要经验是:在生成前减少决策空间可提高准确性和提交的稳健性。我们的代码和支持脚本可在该https URL公开获取。
英文摘要
We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO resolutions. The task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema setting using only open-weight models up to 8B parameters. We frame the task as constrained structured prediction. The system first narrows the candidate tag space with metadata-aware dense retrieval, then applies constrained decoding with per-dimension caps, escalates only uncertain cases to a three-agent debate branch, and finally validates the output schema. On the official leaderboard, LLM-INSTRUCT ranked 1st overall, with 1st in F1 and 5th in LLM-as-a-Judge. During development, our configuration search further improved Task 1b Micro-F1 from 35.83% to 40.08% while keeping the internal Task 2 score at 4.421. The main lesson is simple: reducing the decision space before generation improves both accuracy and submission robustness. Our code and supporting scripts are publicly available at: https://github.com/LLM-Instruct-at-UZH-Shared-Task-2026/Method
发表机构
- Vietnamese-German University(越南-德国大学)
- RMIT University Vietnam(皇家墨尔本理工大学越南分校)
- VANGIA INNOVATIONS(VANGIA创新公司)
机构由 AI 辅助整理,请以论文原文为准。