发表机构
NII LLMC, Japan; National Institute of Informatics, Japan; University of Göttingen, Germany; University of Bologna, Italy; The University of Tokyo, Japan; Inria, LS2N, Nantes Université, France(国立情报学研究所 LLMC,日本; 国立情报学研究所,日本; 哥廷根大学,德国; 博洛尼亚大学,意大利; 东京大学,日本; 法国国家信息与自动化研究所,LS2N,南特大学,法国)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究LLM重写对多模态声明验证的影响,发现模型对自然重写和受控注入具有鲁棒性,但概率偏移存在,且对冲条件影响最大。
AI 中文摘要
已知LLM会在生成的文本中引入风格变化,然而这些风格变化如何影响模型在科学任务上的决策仍未被充分探索。本文聚焦于多模态声明验证,其目标是确定文本声明是否基于给定的证据。我们应用两种重写策略:自然重写,模拟研究人员日常使用LLM润色学术文本的方式;以及受控注入,插入单个与LLM关联的单词以隔离词汇选择的影响。我们评估了11个开放权重模型,涵盖五个VLM系列,参数范围从2B到38B。我们发现模型对这些修改具有鲁棒性:大多数模型的准确率没有显著下降,并且与先前关于评分操纵的研究相比,验证似乎稳定得多。然而,一致的概率偏移确实会发生。对冲导向的条件在几乎所有模型中产生显著偏移,而增强条件显示出较弱的效果,一般润色条件(如语法纠正、流畅性改进)几乎没有影响。
英文摘要
LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.
CommentsAccepted to AACL 2026 (Main Conference). 18 pages