发表机构
Sirjan University of Technology; Department of Computer Engineering(锡尔詹理工大学; 计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出可掩码性指数(MI),用于评估知识关系适合的提示方式,通过掩码与未掩码模板的DepthRank分数差异计算。在ATOMIC2020基准上评估发现MI与下游生成性能正相关,有助于选择提示模板和策略提取预训练语言模型的关系知识。
AI 中文摘要
大规模预训练语言模型如T5和BERT在生成结构化知识方面展现出强大能力,但其性能取决于提示策略与预训练目标的匹配程度。我们引入可掩码性指数(MI),这是一种定量指标,用于估计在少样本生成中知识关系更适合掩码式提示还是前缀式提示。MI通过掩码和未掩码模板之间的DepthRank分数差异计算得出,可衡量目标与模板的对齐程度。我们在ATOMIC2020知识库完成基准的各种关系上评估MI,结果表明它与下游生成性能呈正相关。这些结果表明MI可帮助选择合适的提示模板和适应策略,从预训练语言模型中提取关系知识,特别是在低资源环境中。
英文摘要
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation. MI is computed from differences in DepthRank scores between masked and unmasked templates, providing a principled measure of objective-template alignment. We evaluate MI on a diverse set of relations from the ATOMIC2020 knowledge base completion benchmark and show that it is positively correlated with downstream generation performance. These results indicate that MI can help select appropriate prompting templates and adaptation strategies for extracting relational knowledge from pretrained language models, especially in low-resource settings.