发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出ECCBench基准与评估协议,从效率、压缩性、校准性三方面评估视觉-语言模型的记忆能力,发现预训练VLM对文本记忆可压缩但视频不行,部分非Transformer架构的压缩-校准权衡优于RoPE Transformer。
AI 中文摘要
记忆被广泛视为大语言模型(LLM)与视觉-语言模型(VLM)尚未解决的重要问题,当前基准通常通过测试长文本或视频的准确率来评估记忆,但仅靠准确率会遗漏实际长程任务中重要的特性。我们推出ECCBench,这是一个基准与评估协议,用于测量超越系统容量的记忆能力——即特定预算下的原始准确率——通过三个称为ECC的维度:效率(从记忆中作答所需的浮点运算量FLOPs)、压缩性(是否可压缩输入能被更准确或更高效地记住)、校准性(系统是否会因自身不确定性及错误成本而弃权(不执行))。我们发现,预训练VLM对文本记忆可压缩,但对视频记忆不可压缩,且在两种模态上的校准性均较差。在更广泛的记忆骨干中,若干非Transformer架构比采用旋转位置编码(RoPE)的Transformer实现了更好的压缩-校准权衡,表明它们可能是长程智能体的有用组件。
英文摘要
Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real long-horizon tasks. We introduce ECCBench, a benchmark and evaluation protocol that measures memory beyond a system's capacity--its raw accuracy at a specific budget--via three axes we call ECC: efficiency--the computation, in FLOPs, needed to answer from memory; compression--whether compressible inputs are remembered more accurately or efficiently; and calibration--whether the system abstains in response to its own uncertainty and the cost of an error. We find that pretrained VLMs compress their memory over text but not video and are poorly calibrated on both. Among a broader set of memory backbones, several non-Transformer architectures achieve better compression-calibration tradeoffs than RoPE Transformers, suggesting they may be useful components for agents operating over long horizons.