arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10612cs.CRcs.LG

PEARL:一个用于评估差分隐私合成教育数据的任务感知框架

PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data

Xianghui Meng, Yujing Zhang, Jionghao Lin

首次发表
浏览论文内容

中文总结 AI 辅助

PEARL框架通过多维度检查评估差分隐私合成教育数据,发现多数数据集因遗漏关键组或顺序问题被拒,且隐私保护不保证任务有用性。

中文摘要 AI 辅助

个性化学习系统依赖于真实的学习者数据,包括表现、行为和人口统计信息,但这些数据具有高度隐私敏感性。差分隐私(DP)合成数据可以在减少个体学习者暴露的同时,支持系统开发和教育教学研究。然而,现有的评估方法分别评估隐私性和预测有用性,而没有确定合成学习者数据是否仍可用于预期的个性化学习任务。我们引入了PEARL(隐私等价审计与发布台账),该框架仅当DP合成教育数据集通过有效性、隐私保护、预测有用性和对预期教育任务适用性的所有必要检查时,才批准该数据集,同时记录每个被拒绝数据集失败的原因。在96个研究设置中,每个设置由数据集、数据生成方法、隐私预算和随机种子定义,只有12个产生了通过所有适用PEARL检查的合成数据集。许多受隐私保护的数据集因遗漏重要结果组(如退学学生)或未能保持学习活动的顺序而被拒绝。公平性分析进一步表明,一些通过隐私和预测有用性检查的数据集,在残疾和社会经济背景定义的群体之间,仍产生了不平等的风险预测表现。此外,深度知识追踪和自注意力知识追踪未能从任何测试的合成知识追踪数据集中学习到有意义的下一响应模式,这表明仅靠隐私保护并不能保证对辍学预测、知识追踪或自适应辅导的有用性。

英文摘要

Personalized learning systems rely on real learner data, including performance, behavior, and demographic information, but these data are highly privacy-sensitive. Differentially private (DP) synthetic data can support system development and educational research while reducing exposure of individual learners. Existing evaluations, however, assess privacy and predictive usefulness separately, without determining whether synthetic learner data remain usable for the intended personalized learning task. We introduce PEARL (Privacy-Equivalence Audit and Release Ledger), which approves a DP synthetic educational dataset only when it passes all required checks of validity, privacy protection, predictive usefulness, and suitability for the intended educational task, while recording why each rejected dataset fails. Across 96 study settings, each defined by a dataset, data-generation method, privacy budget, and random seed, only 12 produced synthetic datasets that passed all applicable PEARL checks. Many privacy-protected datasets were rejected for omitting important outcome groups, such as withdrawn students, or for failing to preserve the order of learning activities. Fairness analysis further showed that some datasets passing privacy and predictive-usefulness checks still yielded unequal at-risk prediction performance across groups defined by disability and socioeconomic background. Moreover, Deep Knowledge Tracing and Self-Attentive Knowledge Tracing learned no meaningful next-response patterns from any tested synthetic knowledge-tracing dataset, showing that privacy protection alone does not guarantee usefulness for dropout prediction, knowledge tracing, or adaptive tutoring.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑