arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05655cs.IR

个性化模态加权真的是个性化的吗?多模态推荐器中每用户加权声明的受控审计

Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders

Jingyuan Zheng, Xin Zhang, Yang Gu, Dongjing Wang, Yuxiang Wang, Xudong Shen, Haiping Zhang, Youhuizi Li, Dongjin Yu

AI总结:

该研究审计多模态推荐器的每用户模态加权是否真正个性化,发现单一全局模态权重已提供几乎全部内容增益,每用户加权无一致效用,建议将real-GM与real-shuf作为个性化声明的最低证据标准。

AI中文摘要:

每用户模态加权通过用户模态强度向量、注意力门、元权重超网络和低秩引导权重,在拥有数十亿用户的多模态推荐器中部署,每种方法都声称能从用户特定的模态偏好中获得排名提升。然而,据我们所知,先前的评估并未将真正的用户特定信号与全局模态权重及模型容量隔离开来。我们采用双对比审计原则对该系列方法进行审计,将六种实现方式缩减至一个共享的协同 backbone,并测量相对于单一全局模态权重的效用差距(real-GM),以及相对于用户-权重绑定的评估时置换的可识别性差距(real-shuf)。在三个独立的短视频语料库中,单一全局权重已能提供几乎所有的内容增益(相对于无模态基线,分别提升1.9/3.6/3.5个百分点,p < 0.001)。将权重设为每用户并不会带来一致的效用:没有任何一种实现方式在所有语料库和指标上获胜,少数正差距很小(≤0.9个百分点)且方向不定。置换控制是必要但不充分的,因为对于那些同时比全局权重表现差的头部,real-shuf达到了内容增益的128%。我们将这种分离追溯至门控读取共享协同嵌入:解耦门控输入会将膨胀的real-shuf降至接近零,而效用结论保持不变。单调信号植入剂量反应(捕获AUROC从0.57升至0.89,以及从0.64升至1.00)验证了若存在用户特定结构,该框架将能检测到,且所有发现在第四个跨领域电商语料库上可复现。我们建议将real-GM与real-shuf一同报告,作为个性化声明的最低证据标准。

英文摘要:

Per-user modality weighting is deployed at billion-user scale in multimodal recommenders, through user modality-strength vectors, attention gates, meta-weight hypernetworks, and low-rank guided weights, each claiming a ranking gain from user-specific modality preference. Yet, to our knowledge, prior evaluations do not isolate a genuinely user-specific signal from a global modality weight plus model capacity. We audit this family with a two-contrast audit principle, reducing six implementations onto one shared collaborative backbone and measuring a utility gap (real-GM) against a single global modality weight and an identifiability gap (real-shuf) against an eval-time permutation of the user-weight binding. Across three independent short-video corpora, a single global weight already delivers nearly all of the content gain (+1.9/+3.6/+3.5pp over a no-modality baseline, p < .001). Making the weight per-user adds no consistent utility: no implementation wins on all corpora and metrics, and the few positive gaps are small (<=0.9pp) and flip. The shuffle control is necessary but not sufficient, since real-shuf reaches +128% of the content gain for heads that simultaneously lose to the global weight. We trace this dissociation to gates reading the shared collaborative embedding: decoupling the gate input collapses the inflated real-shuf to near zero while the utility conclusion stands. A monotone signal-implant dose-response (capture AUROC rising from 0.57 to 0.89 and from 0.64 to 1.00) verifies the harness would detect user-specific structure if present, and every finding replicates on a fourth, cross-domain e-commerce corpus. We propose reporting real-GM alongside real-shuf as a minimum evidentiary standard for personalization claims.

↑