发表机构
Zhengzhou University; Harbin Institute of Technology (Shenzhen)(郑州大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RAME框架,结合检索增强少样本选择、三提示集成和多数投票,在无需训练的情况下从育种文献中联合抽取实体和关系,超越GPT-5.5基线11.4%。
AI 中文摘要
本文介绍了我们为CCL2026-Eval任务5:小宗作物育种信息抽取(MGBIE)所开发的系统,该系统从小宗作物育种文献中联合抽取12种实体类型和6种关系类型。我们提出了RAME(检索增强的多提示集成),这是一个无需训练的框架,在受控多样性下激发多个大语言模型输出,并通过多数投票进行聚合以获得高置信度的预测。RAME结合了(i)通过混合BM25-嵌入检索器进行的检索增强少样本选择,(ii)涵盖精确率到召回率谱系的三提示集成(严格、宽松、平衡),以及(iii)大规模重复采样与多数投票以过滤噪声预测。基于DeepSeek-V4-Flash构建的RAME在排行榜上取得了0.499的总分(命名实体识别0.730,关系抽取0.346),排名第一,超过了由GPT-5.5驱动的官方Track-A基线(0.448),相对提升了11.4%。代码可在该https URL获取。
英文摘要
This paper presents our system for CCL2026-Eval Task 5: Minor-Grain Breeding Information Extraction (MGBIE), which jointly extracts 12 entity types and 6 relation types from minor-grain breeding literature. We propose RAME (Retrieval-Augmented Multi-Prompt Ensemble), a training-free framework that elicits multiple LLM outputs under controlled diversity and aggregates them by majority voting to obtain high-confidence predictions. RAME combines (i) retrieval-augmented few-shot selection via a hybrid BM25-embedding retriever, (ii) a three-prompt ensemble (Strict, Relaxed, Balanced) spanning the precision to recall spectrum, and (iii) large-scale repeated sampling with majority voting to filter noisy predictions. Built on DeepSeek-V4-Flash, RAME achieves a Total Score of 0.499 (NER 0.730, RE 0.346) on the leaderboard, ranking 1st and surpassing the official Track-A baseline powered by GPT-5.5 (0.448), representing an 11.4% relative improvement. Code is available at https://github.com/king-wang123/CCL26-RAME.
CommentsAccepted at CCL 2026 (China National Conference on Computational Linguistics)