设计上用于身份识别,偶然泄露人口统计信息:行为生物特征嵌入中的人口统计信息泄露与抑制
Identity by Design, Demographics by Accident: Demographic Leakage and Suppression in Behavioral Biometric Embeddings
浏览论文内容
中文总结 AI 辅助
本文审计行为生物特征认证系统的人口统计信息泄露,评估11种模型、9个数据集及4种抑制方法,发现不同模态泄露与可抑制性差异,揭示该领域隐私风险并提供相关影响因素见解。
中文摘要 AI 辅助
行为生物特征认证(BBA)系统使用深度学习模型将眼动、语音、击键/触笔动态、步态等生物信号转换为用于用户认证的身份嵌入。这些嵌入本应编码身份信息,但可能意外泄露性别、年龄、身高这类敏感人口统计属性。因此,能访问认证模型的攻击者可从生物信号中推断人口统计信息,包括训练或注册时未见过的用户的信息。本文首次对BBA系统中的人口统计信息泄露进行系统审计,评估了9个数据集上的11种模型,涵盖4种生物特征模态;还对Incremental Variable Elimination(IVE)、Hilbert-Schmidt Independence Criterion(HSIC)、Adversarial Encoder-Decoder(AED)、Protected Attribute Suppression System(PASS)这4种事后抑制方法进行基准测试,以评估它们在减轻人口统计信息泄露的同时保留认证效用的能力。分析发现,不同模态、模型架构和学习目标间的泄露程度与可抑制性存在显著差异:语音嵌入泄露程度高且可被有效抑制,而击键/触笔嵌入泄露程度较低但难以净化。这些发现凸显了行为生物特征认证中的一项基本隐私风险,并为人口统计信息可抑制性的影响因素提供了见解。
英文摘要
Behavioral biometric authentication (BBA) systems use deep learning models to transform biometric signals, such as eye movements, voice, keystroke/touchstroke dynamics, and gait, into identity embeddings for user authentication. While designed to encode identity, these embeddings may inadvertently reveal sensitive demographic attributes, including gender, age, and height. Consequently, an adversary with access to the authentication model can infer demographic information from biometric signals, including those of users unseen during training or enrollment. In this paper, we present the first systematic audit of demographic leakage in BBA systems, evaluating 11 models across 9 datasets spanning four biometric modalities. We further benchmark four post-hoc suppression methods-Incremental Variable Elimination (IVE), Hilbert-Schmidt Independence Criterion (HSIC), Adversarial Encoder-Decoder (AED), and Protected Attribute Suppression System (PASS) to assess their ability to mitigate demographic leakage while preserving authentication utility. Our analysis reveals substantial variation in leakage and suppressibility across modalities, model architectures, and learning objectives. While voice embeddings exhibit high leakage that can be effectively suppressed, keystroke/touchstroke embeddings exhibit lower leakage but are considerably more difficult to sanitize. These findings highlight a fundamental privacy risk in behavioral biometric authentication and provide insights into the factors governing demographic information suppressibility.
发表机构
- Department of Computer Science and Engineering, University of Moratuwa(莫拉图瓦大学计算机科学与工程系)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。