发表机构
Truth Audit Labs(真相审计实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将识别未知大语言模型的问题形式化为外部探针设计与内部序贯识别两层,外部问题等价于加权集合覆盖,内部问题给出评估次数上界。
AI 中文摘要
主动序贯假设检验研究如何利用一组给定的感知动作来识别未知假设。我们在识别大型语言模型(LLMs)的场景中研究这一问题,即,如果用户正在与从一组已知模型中抽取的一个LLM进行对话,他们如何识别当前使用的是哪一个模型?在此,可用的感知动作(评估)本身就是一个设计选择:评估者必须首先决定构建哪些环境和提示族,然后才决定如何序贯地使用它们。我们将这两个层面形式化为一个外部探针设计问题和一个内部识别问题。简言之,外部阶段选择一组探针发送给整个模型集合,从而创建一种指纹数据集。随后是内部阶段,该阶段序贯地发送一组预算最小化的探针以识别正在使用的模型。对于外部问题,我们证明,以最小成本选择要构建的评估,使得每对候选模型都能被区分,这恰好是一个加权集合覆盖问题。由于候选模型的响应分布并非精确已知,而仅通过校准样本获得,我们给出了一种一次性程序,从这些样本中估计覆盖实例。对于内部问题,我们根据可用评估区分每对候选模型的程度,界定了识别未知模型所需的评估次数。
英文摘要
Active sequential hypothesis testing studies how to identify an unknown hypothesis with a given set of sensing actions. We study this in the setting of identifying large language models (LLMs), \textit{i.e.}, if a user is conversing with an LLM drawn from a known set of models, how can they identify which one is in use? Here, the available sensing actions (evaluations) are themselves a design choice: an evaluator must first decide which environments and prompt families to construct, and only then decide how to use them sequentially. We formalize these two levels as an outer probe-design problem and an inner identification problem. Simply put, the outer stage selects a set of probes to be sent to the entire set of models, creating a kind of fingerprint dataset. This is followed by the inner stage, which sequentially sends a budget-minimizing set of those probes to identify the model in use. For the outer problem, we show that selecting which evaluations to construct at minimum cost, so that every pair of candidates is distinguished, is exactly a weighted set cover problem. Since the response distributions of the candidate models are not known exactly but only through calibration samples, we give a one-shot procedure that estimates the cover instance from these samples. For the inner problem, we bound the number of evaluations needed to identify the unknown model in terms of how well the available evaluations distinguish each pair of candidates.