arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39846cs.CLcs.AIcs.HC

当幼儿园学生解微积分:测量角色提示推理模型中的能力泄漏

When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models

Pakhapoom Sarapat, Saksorn Ruangtanusak, Kunat Pipatanakul, Pittawat Taveekitworachai

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出RoleCapBench基准,发现角色提示模型存在能力泄漏,并设计Injection干预方法,有效降低超角色准确率,提升角色能力对齐。

中文摘要 AI 辅助

我们研究了角色能力泄漏(RCL)问题,即角色提示推理模型在生成令人信服的角色内文本的同时,在基准测试中持续展现出超出所分配角色隐含能力的能力。例如,当模型被提示扮演幼儿园学生的角色时,人们可能期望其在数学基准上的表现反映幼儿园水平的能力,而非解决微积分问题的专家级熟练度。我们引入了RoleCapBench,一个基于课程体系的基准,用于评估从小学到A-level的六个教育角色和四个评估水平上的RCL,并使用它评估三个开放权重推理模型。我们发现,尽管模型能够生成风格上令人信服的角色内响应,但它们始终未能将其潜在能力与所分配的角色对齐。简单的角色提示产生了1.218至1.389的强角色语音得分,同时保持了0.811至0.898的超角色准确率。RCL在一系列提示条件下持续存在,包括明确指示模型匹配角色能力水平的提示。为缓解此问题,我们提出了Injection,一种推理时干预方法,结合了明确的、特定角色的能力指南和引导性的预填充响应前缀。Injection改善了模型间的角色能力对齐,将超角色准确率降低了多达0.562,同时在大多数模型上保持了角色内准确率,仅出现小于0.058的微小下降。所有工件,包括脚本和评估数据,将在录用后发布。

英文摘要

We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level proficiency in solving calculus problems. We introduce RoleCapBench, a curriculum-grounded benchmark for evaluating RCL across six educational roles and four assessment levels spanning elementary school through A-level, and use it to evaluate three open-weight reasoning models. We find that although the models can generate stylistically convincing in-role responses, they consistently fail to align their underlying capabilities with their assigned roles. Naive role prompting yields strong role-voice scores of 1.218--1.389 while retaining above-role accuracy of 0.811--0.898. RCL persists across a range of prompting conditions, including prompts that explicitly instruct the model to match the role's capability level. To mitigate this problem, we propose Injection, an inference-time intervention that combines explicit, role-specific capability guidelines with a guiding prefilled response prefix. Injection improves role-capability alignment across models, reducing above-role accuracy by up to 0.562 while preserving in-role accuracy with a marginal drop of less than 0.058 across most models. All artifacts, including scripts and evaluation data, will be released upon acceptance.

发表机构

  • SCB DataX
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑