发表机构
Beijing University of Posts and Telecommunications; Tianjin University(北京邮电大学; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出从单一基础模型表示预测多模型间的表示分歧(拉什蒙表示),验证其输入依赖性与可预测性,并训练轻量级预测器实现高效可靠性估计。
AI 中文摘要
基础模型正日益广泛地应用于各类场景,通常作为AI系统的核心模块。然而,不同的基础模型可能从多个不同视角编码同一输入,导致显著的表示分歧,我们将其称为“拉什蒙表示”。这种分歧通常表明,给定模型对某输入的编码方式与其他模型不一致,为输入可靠性估计提供了一个有价值但尚未充分探索的信号。尽管先前的工作主要集中于通过表示集合来测量多个模型之间的分歧,我们则专注于从单一表示预测分歧。我们假设这种分歧遵循某种一致的、依赖于输入的模式,而非随机发生。为验证这一点,我们通过比较每个样本在不同模型表示空间中的最近邻来量化分歧,然后训练一个轻量级预测器,从单一模型的表示估计分歧。在推理时,给定新输入,预测器利用该输入的表示来判断其与其他模型的表示是否一致或偏离。跨多种基础模型和数据集的广泛实验表明,表示分歧确实依赖于输入、可预测且可泛化,从而能够实现基础模型的高效可靠性估计。
英文摘要
Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundation models may encode the same input from multiple different views, leading to substantial representation disagreement, which we term Rashomon Representation. Such disagreement often signals inputs that a given model encodes in a way inconsistent with other models, offering a valuable yet underexplored signal for input reliability estimation. While prior work has largely focused on measuring disagreement across multiple models with a representation set, we instead focus on predicting disagreement from a single representation. We hypothesize that this disagreement follows some consistent, input-dependent patterns rather than occurring at random. To test this, we quantify disagreement by comparing each sample's nearest neighbors across different models' representation spaces, then train a lightweight predictor that estimates disagreement from a single model's representation. At inference time, given a new input, the predictor uses that input's representation to tell whether it aligns with or diverges from those of other models. Extensive experiments across diverse foundation models and datasets show that representational disagreement is indeed input-dependent, predictable, and generalizable, enabling efficient reliability estimation of foundation models.
CommentsUnder review