发表机构
Academy of Cryptography Techniques; Government Cipher Committee; Ho Chi Minh City University of Technology (HCMUT); Vietnam National University Ho Chi Minh City(密码技术学院; 政府密码委员会; 胡志明市理工大学; 越南胡志明市国家大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对越南RoPA合规需求,提出RoPA Manager系统,结合混合检索与本地LLM自动提取处理活动记录,构建基准并验证了本地模型与云模型性能相当。
AI 中文摘要
越南《个人数据保护法》(第91/2025/QH15号法律)和第356/2025/ND-CP号法令,自2026年1月1日起生效,要求组织建立并维护处理活动记录(RoPA)。手动编制RoPA劳动密集,而云端托管的大语言模型(LLM)可能与数据主权要求相冲突。我们提出了RoPA Manager,一个使用混合检索的自动化RoPA信息提取系统,该系统结合了基于tsvector的词法排序、稠密向量搜索、倒数排名融合(RRF)以及本地部署的LLM。我们引入了一个越南语RoPA基准,包含32个组织、77个处理活动、12个字段组和4,338个参考值。评估在三个不同层面进行。自动化评分器在扰动数据上测试(不调用LLM),实现了F1=0.9493 [0.9436, 0.9548];这衡量的是评分器的稳健性而非端到端提取的准确性。端到端提取对参考标签的token覆盖率为50.04%-55.25%。两位独立专家审查了1,558个参考值(占基准的35.9%),未发现错误值,并实现了99.68%的一致性,PABAK=0.9936。未测量值级精确度。在24 GB GPU上的32个配对场景中,本地部署的Qwen3.5-27B-GPTQ-Int4与基于云的DeepSeek-V4-Flash相比无统计学显著差异(差异0.20个百分点,有利于DeepSeek,95% CI [-0.93, 1.32],p=0.72),而Gemma-4-31B表现显著更差(p<0.01)。
英文摘要
Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs) may conflict with data-sovereignty requirements. We propose RoPA Manager, a system for automated RoPA information extraction using hybrid retrieval that combines lexical ranking over tsvector, dense-vector search, Reciprocal Rank Fusion (RRF), and locally deployed LLMs. We introduce a Vietnamese RoPA benchmark with 32 organizations, 77 processing activities, 12 field groups, and 4,338 reference values. Evaluation is reported at three distinct levels. The automated scorer, tested on perturbed data without invoking an LLM, achieved F1 = 0.9493 [0.9436, 0.9548]; this measures scorer robustness rather than end-to-end extraction accuracy. End-to-end extraction achieved token coverage of 50.04-55.25% against the reference labels. Two independent experts reviewed 1,558 reference values (35.9% of the benchmark), found no incorrect values, and achieved 99.68% agreement with PABAK = 0.9936. Value-level precision was not measured. Across 32 paired scenarios on a 24 GB GPU, locally deployed Qwen3.5-27B-GPTQ-Int4 showed no statistically significant difference from cloud-based DeepSeek-V4-Flash (difference 0.20 percentage points in favor of DeepSeek, 95% CI [-0.93, 1.32], p = 0.72), while Gemma-4-31B performed significantly worse (p < 0.01).
CommentsEnglish version followed by Vietnamese version. Accepted for publication in the Proceedings of the 29th National Conference on Selected Issues of Information and Communication Technology (VNICT 2026), Hanoi, Vietnam, November 7-8, 2026