arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HABIB_TAZ参加SemEval-2026任务11:通过合成训练和多目标优化从内容中分离形式逻辑

HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

Abdullah Shaikh, Zain Naqi, Taha Zahid, Sandesh Kumar, Abdul Samad

arXiv 2607.14349首次发表:更新:

发表机构

Dhanani School of Science & Engineering Habib University(哈比卜大学达纳尼科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对SemEval-2026任务11,用基于三段论合成数据集微调的mDeBERTa-v3网络,结合多目标损失函数应对大语言模型形式推理受内容影响的问题,在多个子任务中取得优异成绩,还公开了数据集生成引擎和代码库。

AI 中文摘要

虽然大语言模型在许多通用自然语言处理任务中表现出色,但其形式推理能力常受内容影响而受损。本文介绍了参加SemEval-2026任务11的系统,该任务评估模型在有和无干扰前提的情况下跨12种语言从内容中分离形式逻辑的能力。我们使用在基于规则的三段论合成数据集上微调的mDeBERTa-v3网络应对挑战。训练管道采用多目标损失函数,结合自适应组分布鲁棒优化、可调度的可微偏差惩罚和KL散度一致性正则化。在多个子任务中取得了优异成绩,数据集生成引擎和代码库公开可用。

英文摘要

While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we present our system for SemEval-2026 Task 11, which evaluates the ability of models to disentangle formal logic from content across 12 languages with and without distractor premises. We address this challenge using mDeBERTa-v3 networks fine-tuned on a synthetic, rule-based dataset of syllogistic schemes to avoid the semantic noise of LLM-augmented data. To explicitly decouple plausibility from logical structure, our training pipeline employs a multi-objective loss function combining Adaptive Group Distributionally Robust Optimization (DRO), a scheduled differentiable bias penalty, and KL-Divergence consistency regularization. Our system achieved #1 ranks and perfect Ranking Scores (100.0) with 0.00% bias and 100.0% accuracy on Subtask 1 (English), Subtask 2 (Noisy English), and Subtask 3 (Multilingual). On the highly complex Subtask 4 (Noisy Multilingual), the system achieved the 6th rank with 89.06% Accuracy and F1-score, alongside a limited 2.89% Bias and a 37.78 Ranking Score. Our dataset generation engine and codebase are publicly available to facilitate future work on robust logical reasoning.

Journal refProceedings of the 20th International Workshop on Semantic Evaluation (2026), pp. 1006-1014 (2026)

DOI:10.18653/v1/2026.semeval-1.139

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑