arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11816cs.CRcs.AIcs.CL

中国起源的多模态视觉语言模型如何在状态对齐中从拒绝转向重构

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of California, Berkeley(加利福尼亚大学伯克利分校)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian

中文总结 AI 辅助

该研究构建基准测试9个视觉语言模型,发现中国起源模型更易进行状态对齐重构,且审查从可见的拒绝转向不可见的重构,这对人机交互存在影响。

中文摘要 AI 辅助

已有研究记录了中国起源的基于文本的大型语言模型(LLM)存在状态对齐扭曲问题,但多模态系统中是否存在该问题以及以何种形式存在尚未得到系统研究。我们构建了包含10个政治敏感主题的200个核心条目的平衡基准,以及一个包含7种变体的视觉抽象探针,对9个视觉语言模型(VLM)进行测试,其中7个为中国起源,2个为非中国起源,测试涵盖4种提示范式和2种提示语言,共产生21708次试验。每个响应由两名独立的前沿LLM法官在6个维度上进行审核,包括明确拒绝、信息完整性、视觉接地、状态对齐重构、语言一致性和响应长度,并对200次试验样本与3名人类专家进行验证。通过分别测量每个维度,我们将多模态审查分解为单独的信号,而非单一的基于拒绝的评分;特别是,拒绝和重构是独立测量的,因此模型可以停止拒绝但仍进行重构。我们发现:(i)在所有模型中,中文提示使状态对齐重构的几率大约增加了两倍;(ii)中国起源模型的重构程度高于非中国起源模型(该方向在法官和人类评估者中均稳定,幅度为1.6至3.2倍);(iii)该效应在纯文本政治评论中最强(36.5%),且由对所描绘主体的识别而非像素细节决定,即使是标志性图像的轮廓也存在该效应;(iv)在4代Qwen模型中,状态对齐重构上升而明确拒绝下降:审查从可见行为(拒绝)转向不可见行为(流畅的重构)。我们认为,这种向不可见重构的转变本质上是人机交互的问题:它消除了用户用来识别信息被扣留的信号。

英文摘要

State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.

补充信息

↑