AI 中文总结
本研究提出CARGO-VL框架,结合XMC资源,通过组相对优化及原对偶控制器,提升视觉-语言模型的反事实仲裁性能,在多基准上优于逐点基线。
AI 中文摘要
视觉-语言系统将图像与检索文本结合,但这些来源可能意见不一致或共同无法支持答案。可靠的模型必须识别可信来源,并在两者都不足时弃权(不执行)。现有的后训练目标独立对实例打分,因此无法在反事实证据变化下执行连贯行为。我们提出CARGO-VL,一种组相对框架,该框架将涵盖对齐、图像正确、文本正确和两者错误(A/V/T/N)证据状态的匹配变体作为一个整体优化。其目标将条件正确性与答案不变性、来源等变性及答案到弃权(不执行)切换的转换奖励相结合,同时原对偶控制器平衡不安全答案与过度 deferral。我们还提供XMC(扩展模态冲突),一种四条件冲突训练资源,并在CMC-Bench和Modality-Bench上评估迁移性能。在多个随机种子下,CARGO-VL在冲突处理、无支持答案避免和模态平衡方面优于逐点基线。消融实验确定了关系转换信号和自适应风险控制的互补益处,支持反事实一致性作为可靠多模态证据仲裁的实用目标。
英文摘要
Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing post-training objectives score instances independently and therefore do not enforce coherent behavior under counterfactual evidence changes. We introduce CARGO-VL, a group-relative framework that optimizes matched variants covering aligned, image-correct, text-correct, and both-wrong (A/V/T/N) evidence states as one bundle. Its objective couples condition-wise correctness with transition rewards for answer invariance, source equivariance, and answer-to-abstention switching, while a primal-dual controller balances unsafe answers against excessive deferral. We also contribute XMC (eXtended Modal Conflict), a four-condition conflict training resource, and evaluate transfer on CMC-Bench and Modality-Bias. Across multiple seeds, CARGO-VL improves conflict handling, unsupported-answer avoidance, and modality balance over pointwise baselines. Ablations identify complementary benefits from relational transition signals and adaptive risk control, supporting counterfactual consistency as a practical objective for reliable multimodal evidence arbitration.