AI 中文总结
该研究提出Predictive Memory Localization(PML),通过分离随机校准目标移动与语义邻居损伤等,在多数据集实验中验证其可实现选择性干预路径的预测与风险感知决策。
AI 中文摘要
激活引导将局部表征转化为控制方向,但仅定位无法揭示该方向是否具有选择性运行机制。我们提出预测性记忆定位(Predictive Memory Localization, PML),将测量网格干预路径作为记忆定位的预测对象。PML将随机校准的目标移动与语义邻居及能力损伤分离,通过强度不相交的低剂量因果响应对比静态定位与监督几何。我们的冻结研究涵盖9个数据集、14个领域的3000条记录,产生30000条不同的记录-方向-层路径及210000条不同的路径-强度评估。在第7层,几何推导的RFM/AGOP方向达到13.1%的目标任意(target-any)和12.3%的干净任意(clean-any),在记录配对自助法下较随机分别高出3.6和3.4个百分点。在按记录、数据集和领域分组的划分中,|α|=0.1时的响应是不相交强度|α|∈{0.25,0.5}下结果的最强信号。在保留记录上,预测器驱动的选择器选择系数或弃权(不执行),相较于训练调整的固定强度策略,提升效用并减少语义邻居损伤,且避免密集扫描中的大部分评估。在3个残差范数匹配的基础模型上,学习到的方向保留选择性路径增益,低剂量响应产生0.801-0.828的记录保留宏平均AUROC。因此,PML将记忆定位转化为边际级选择性结果的可证伪预测及风险感知干预决策。
英文摘要
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|α|=0.1$ are the strongest signal for outcomes at disjoint strengths $|α|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.