arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

这是谁的记录?个性化多模态模型中的记录使用诊断与授权

Whose record is this? Diagnosing and authorizing record use in personalized multimodal models

Xinyu Mao, Junsi Li, Chenyang Liu, Haoji Zhang, Ming Sun

arXiv 2609.04801首次发表:更新:

发表机构

University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对个性化多模态模型的视觉记忆错绑问题,提出记录授权框架,构建RecordAuth-Diag诊断套件,验证了类型化预生成授权可降低未授权记录使用,揭示决策由答案支持主导。

AI 中文摘要

上下文视觉个性化可检索到真实记录,但可能将其应用于错误的视觉主体。我们将记录可用于为答案提供条件的情况形式化为「记录授权」:主体存在(P)、记录-边有效性(E)和答案支持(S)必须同时成立。我们将此类违反称为视觉记忆错绑(VMM)。我们构建了RecordAuth-Diag,这是一个包含3690个案例的匹配诊断套件,在保持查询、问题、记录文本和图像集合固定的情况下,改变一个图像-记录边。移除卡片和随机数重标记将这些失败归因于提供的记录。Raw-bank失败涉及Qwen、Phi和Gemma系列接口:Gemma-3-4B-IT在25.75%的干净召回率下达到63.69%的本地未授权使用。CoViP在类似的干净召回率下为26.02%,而其Qwen骨干网络为22.49%。类型化预生成授权将Qwen在RecordAuth-Diag上的卡片暴露从43.63%降至3.06%,同时正召回率从86.26%变为60.90%。完整的P∧E∧S验证使用560个本地化DAVIS案例:Top-1相关性和类型化授权的释放率相当(分别为28.93%和28.39%),但不安全释放率分别为6.79%和0.89%。在移除的33个额外不安全案例中,27个属于支持类,4个属于边类,2个属于干净类,0个属于边界类。因此,观察到的增量是由支持主导的E∧S决策,而非单独的边检查。外观仅在P的条件下提供E证据;经认证的主体令牌将缺失的存在见证实例化为充分性控制。这些主张涉及评估的约定,而非自然流行率、同意或视觉身份。

英文摘要

Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We formalize when a record may condition an answer as \emph{record authorization}: subject presence ($P$), record-edge validity ($E$), and answer support ($S$) must all hold. We call violations visual memory misbinding (VMM). We construct RecordAuth-Diag, a 3,690-case matched diagnostic suite that changes one image--record edge while holding the query, question, record text, and image multiset fixed. Card removal and nonce relabeling attribute these failures to supplied records. Raw-bank failures span Qwen-, Phi-, and Gemma-family interfaces: Gemma-3-4B-IT reaches 63.69\% local unauthorized use at 25.75\% clean recall. CoViP remains at 26.02\%, versus 22.49\% for its Qwen backbone at similar clean recall. Typed pre-generation authorization reduces Qwen card exposure on RecordAuth-Diag from 43.63\% to 3.06\%, while positive recall changes from 86.26\% to 60.90\%. Full $P\wedge E\wedge S$ validation uses 560 localized DAVIS cases: top-1 relevance and typed authorization have comparable release (28.93\% and 28.39\%) but 6.79\% and 0.89\% unsafe release, respectively. Of the 33 additional unsafe cases removed, 27 are support, 4 edge, 2 clean, and 0 boundary cases. Thus the observed increment is an $E\wedge S$ decision dominated by support, not an edge check alone. Appearance supplies $E$ evidence only conditional on $P$; authenticated subject tokens instantiate the missing presence witness as a sufficiency control. The claims concern the evaluated contracts, not natural prevalence, consent, or visual identity

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑