ActMap:基于生成时激活图的单次前向不确定性量化
ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps
浏览论文内容
中文总结 AI 辅助
ActMap通过压缩生成时激活图为固定张量,实现单次前向的不确定性量化,轻量分类器高效预测正确性,优于多种基线,支持弃权与路由。
中文摘要 AI 辅助
大型语言模型的实用不确定性量化(UQ)必须从单次生成中决定某个特定答案是否可信。现有方法要么采样多次生成,要么只读取输出词元概率,要么将模型的内部计算压缩为单个隐藏状态。我们引入了ActMap,一种白盒表示,它将生成时的隐藏状态轨迹(每一层、每个生成的词元)压缩为一个固定的$12 \ imes 32 \ imes 128$的时间统计通道张量,该张量保留了跨Transformer深度和池化隐藏坐标的结构。该映射在生成过程中捕获,无额外开销,在模型深度和隐藏大小上具有固定形状,占用96 KiB:这是一种紧凑的工件,可保留用于审计相关的生成,并可直接探测,通过遮挡分析将分类器的信号定位到映射的中深度区域。一个轻量级分类器,实例化为紧凑的视觉Transformer,在不到一毫秒的时间内从每个映射读取估计的正确性概率;容量匹配的MLP表现相当,表明表示本身承载了结果。在短答案问答、直接答案数学和摘要事实性上,使用三个指令微调的7-8B模型进行域内训练和评估,ActMap始终优于采样、词元概率、注意力和嵌入基线,并匹配ACT-ViT(一种在密集激活张量上训练、大小大67倍的检测器),在十个十二对中的平均AUROC基本相同,校准误差更低。所得分数支持从单次生成中进行弃权(不执行)、路由和选择性验证,使其成为可扩展监督部署模型的实际原语。
英文摘要
Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's internal computation to a single hidden state. We introduce ActMap, a white-box representation that compresses the generation-time hidden-state trajectory (every layer, every generated token) into a fixed $12\times32\times128$ tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates. The map is captured during the generation pass with no measurable overhead, has a fixed shape across model depths and hidden sizes, and occupies 96 KiB: a compact artifact that can be retained for audit-relevant generations and probed directly, with occlusion analysis localizing the classifier's signal to mid-depth regions of the map. A lightweight classifier, instantiated as a compact Vision Transformer, reads an estimated correctness probability from each map in a fraction of a millisecond; capacity-matched MLPs perform comparably, indicating the representation itself carries the result. Trained and evaluated in-domain on short-answer QA, direct-answer math, and summarization factuality with three instruction-tuned 7-8B models, ActMap consistently outperforms sampling, token-probability, attention, and embedding baselines, and matches ACT-ViT, a detector trained on dense activation tensors $67\times$ larger, at essentially the same mean AUROC with lower calibration error on ten of twelve pairs. The resulting score supports abstention, routing, and selective verification from a single generation, making it a practical primitive for scalable oversight of deployed models.
发表机构
- University of Bologna(博洛尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。