发表机构
WritersLogic Inc(WritersLogic 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该工作参与 CLEF 2026 SimpleText 任务,提出多候选 LLM 简化流水线与基于 NLI 的复杂度识别方法,在句子级简化及幻觉检测上取得领先性能。
AI 中文摘要
我们描述了 Writerslogic 团队参与 CLEF 2026 SimpleText 共享任务的情况,涉及任务1(文本简化)和任务2(复杂度识别)。对于任务1,我们开发了一个使用 GPT-4o-mini 的多候选生成流水线,该流水线以不同温度为每个句子生成五个简化候选,然后使用一种无参考的评分启发式方法选择最佳候选,该方法奖励压缩、源词保留、Cochrane 简明语言摘要词汇使用以及词汇简洁性。在任务1.1(句子级简化)上,我们的 Claude Sonnet 4 提交取得了 SARI 47.43 和 BLEU 14.21,是排名最高的句子级系统(在任务1综合排行榜上位列第三,仅次于两个文档级提交)。对于任务2,我们在 350K 个带标签的(源,句子)对上微调了一个 DeBERTa-v3-large NLI 模型,将幻觉检测视为自然语言推理。该模型将最相关的源句子作为前提,候选句子作为假设,直接学习区分有根据的内容与幻觉内容。在任务2.1(二元过度生成识别)上,我们微调的 DeBERTa 系统取得了 0.8081 的文档级宏 F1(最佳集成中为 0.8085),在识别赛道中排名第一,在所有团队中总体排名第二,仅次于 AIIR Lab(0.8197)。在任务2.2(多类错误分类)上,我们最好的提交达到了 0.804 的多类准确率,在独立团队中排名第二,仅次于 AIIR Lab(0.827)。我们在来自 Cochrane 系统评价的英语和多语言生物医学文本上评估了这两个任务。
英文摘要
We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 (sentence-level simplification), our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system (3rd on the combined Task 1 leaderboard, behind two document-level submissions). For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled (source, sentence) pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 (binary overgeneration identification), our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 (0.8085 in our best ensemble), the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab (0.8197). On Task 2.2 (multi-class error classification), our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab (0.827). We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.
Comments11 pages, 3 tables. Notebook for the SimpleText Lab at CLEF 2026. Code: https://github.com/dcondrey/simpletext-clef2026
Journal refCLEF 2026 Working Notes, CEUR Workshop Proceedings, pp. 6068-6078