arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13568cs.CLcs.LG

语言模型中的分级实体熟悉度读数:波兰语适配、跨语言鲁棒性和拒绝引导

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering

  • Prosit AS(Prosit公司)

机构由 AI 辅助整理,请以论文原文为准。

Grzegorz Brzezinka

AI总结:

研究语言模型能否在生成答案前估计对实体的熟悉度,通过新数据集及实验发现熟悉度探测分数可区分实体,在波兰语适配家族中追踪实体受欢迎程度,跨语言鲁棒,在特定模型中可调节拒绝率,校准后的熟悉度探测有竞争力。

AI中文摘要:

我们研究了来自 Bielik、PLLuM、Gemma-4 和 Qwen3 家族的 12 个指令调优模型在最终提示令牌处的激活情况,使用了一个包含 1440 个波兰实体的新数据集。熟悉度探测分数在每个家族中都能区分真实实体和虚构实体。在波兰语适配的 Bielik 和 PLLuM 家族中,它们还能追踪实体受欢迎程度。在配对实验中,探测对提示语言具有鲁棒性。在 Gemma-4-12B 中,添加熟悉度方向可调节拒绝率。校准后的熟悉度探测在预生成弃权门中具有竞争力。这些结果支持分级预生成实体熟悉度读数以及表征熟悉度与转换为弃权的策略之间的分离。

英文摘要:

Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve instruction-tuned models from the Bielik, PLLuM, Gemma-4, and Qwen3 families, using a new dataset of 1,440 Polish entities spanning four domains and ten Wikipedia-pageview deciles, plus fabricated controls. Familiarity-probe scores separate real from fabricated entities in every family; in the Polish-adapted Bielik and PLLuM families they additionally track entity popularity (model-mean Spearman $ρ$ 0.28-0.57, versus at most 0.11 in Gemma-4 and Qwen3), a pattern more strongly associated with Polish adaptation than with parameter count in this model sample. In a paired experiment on two families, probes retain 96-101% of within-language AUROC when the Polish question stem is replaced with an English one around unchanged entity names, showing robustness to prompt language in this setting. In Gemma-4-12B, the only model that natively refuses, adding a one-dimensional familiarity direction at a single layer moves refusal rates monotonically in both directions (0.24 to 1.00 on well-known entities; 0.73 to 0.00 on unknown ones). Finally, a calibrated familiarity probe is competitive among pre-generation abstention gates, although post-generation detectors better predict behavioral error on average. These results support a graded pre-generation entity-familiarity readout, and a separation between representational familiarity and the policy that converts it into abstention.

补充信息

↑