arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM辅导器中教学式泄露的可审计发布控制

Auditable Release Control for Pedagogical Leakage in LLM Tutors

Nizam Kadir

arXiv 2608.00515首次发表:更新:

发表机构

Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM辅导器的教学式泄露问题,提出感知授权的可审计发布控制机制,经实验验证可显著降低泄露标记,且在安全与效用上优于基线方法。

AI 中文摘要

大型语言模型(LLM)辅导器可以正确且有帮助,但可能在授权披露前泄露答案或决定性推理。我们将这种依赖状态和动作的故障形式化为教学式泄露,并引入一种感知授权的完全中介边界。选择器会发出五种披露契约之一,可信策略管控特权模式,渲染器生成语言。单个发布函数应用可检查的检查项、可选的累积验证以及特定动作的回退;可重放的轨迹将选择、生成、验证和执行故障分离开来。匹配的组件归因揭示了安全-效用前沿。在599个固定的Gemini 3.5提案上,严格的中介将盲测三模型面板多数的泄露标记从181个降至0个(配对问题集群差异为-30.22个百分点,95%置信区间[-35.00,-25.72]),同时替换了581个响应并降低了帮助性。仅检查器触发的回退产生11个多数标记;添加语义验证器产生14个,无可靠边际增益。全局A1支架产生0个多数标记和54个任意评判者标记,在自动安全和效用方面优于拟合Q。在对40个未见过的问题集群和480个攻击序列的外部时间戳复现中,高保证发布将多数标记从42个降至8个(-7.08个百分点,95%置信区间[-13.13,-2.29]);7个故障持续存在,1个被引入,平均帮助性下降0.192。这些结果确立了在声明契约下的可审计发布边界和故障归因,而非通用语义安全或学习增益。

英文摘要

Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize this state- and action-dependent failure as pedagogical leakage and introduce an authorization-aware complete-mediation boundary. A selector emits one of five disclosure contracts, trusted policy gates privileged modes, and a renderer proposes language. A single release function applies inspectable checks, optional cumulative verification, and action-specific fallback; replayable traces separate selection, generation, verification, and enforcement failures. Matched component attribution exposes a safety-utility frontier. On 599 fixed Gemini 3.5 proposals, strict mediation reduces blinded three-model panel-majority leakage flags from 181 to 0 (paired problem-cluster difference -30.22 points, 95% CI [-35.00,-25.72]), while replacing 581 responses and lowering helpfulness. Checker-triggered fallback alone yields 11 majority flags; adding the semantic verifier yields 14 and no reliable marginal gain. A global A1 scaffold yields 0 majority and 54 any-judge flags, outperforming fitted Q on automatic safety and utility. In an externally timestamped replication over 40 unseen problem clusters and 480 attack sequences, high-assurance release reduces majority flags from 42 to 8 (-7.08 points, 95% CI [-13.13,-2.29]); seven failures persist, one is introduced, and mean helpfulness falls by .192. These results establish an auditable release boundary and failure attribution under declared contracts, not universal semantic safety or learning gains.

Comments9 pages, 1 figure, 6 tables. Preprint; not peer reviewed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑