arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预算受限的具身感知:四道资源墙及对小于31B的开放模型的访问结构化感知的预注册评估

Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

Defu Lin, Wenhui Chen, Ziyao Lin, Jianlin Chen, Peiji Long, Chi Man Vong

arXiv 2608.22975首次发表:更新:

发表机构

University of Macau; South China University of Technology(澳门大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对具身多模态智能体的固定决策token预算问题,提出四道资源墙及无训练包装器ASP,在SEW-Bench上评估发现查询条件访问比参数数量更关键。

AI 中文摘要

具身多模态智能体必须在固定的每决策token预算下,从不断增长的观察流中作出回答。我们通过四道资源墙将这一约束形式化:用于状态受限的感知香农墙、用于与查询无关的帧选择的视界墙、用于非自适应检索的轮次墙,以及用于固定深度推理的条件组合墙。我们引入ASP,这是一种针对冻结多模态模型的无训练包装器,它结合了 capped structured state( capped结构化状态)、逐字情节索引,以及带有迭代访问的查询条件预算分配。按照预注册协议,我们在SEW-Bench(一个为实例化这些资源墙而构建的无许可合成长视界走查基准)上评估了7个3B至31B的开放权重模型。我们未运行已注册的自然视频基准,因为它们的帧需要数据集协议;因此我们的证据涉及访问机制,而非自然场景感知。在4096-token的决策预算下,ASP达到75%至94%的情节检索准确率,而同等预算下的与查询无关的采样仅为3%至19%,且在每个主干模型上,预算重新分配的表现优于将采样预算增加三倍。然而,完整的三组件架构未验证通道对偶性:移除压缩状态使旗舰模型的均值从35.4提升至58.0,ASP在任何主干模型上均未优于仅逐字的基线,且四项预注册证伪标准中有两项触发。这些结果表明,在固定预算下,查询条件访问而非仅参数数量或上下文增长具有决定性,而在此设置中,提示的在线压缩并未体现其成本效益。

英文摘要

Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constraint through four resource walls: a perceptual Shannon wall for bounded state, a horizon wall for query-independent frame selection, a round wall for non-adaptive retrieval, and a conditional composition wall for fixed-depth inference. We introduce ASP, a training-free wrapper for frozen multimodal models that combines a capped structured state, a verbatim episodic index, and query-conditioned budget allocation with iterative access. Following a pre-registered protocol, we evaluate seven open-weight models from 3B to 31B on SEW-Bench, a license-free synthetic long-horizon walkthrough benchmark constructed to instantiate these walls. The registered natural-video benchmarks were not run because their frames require dataset agreements; our evidence therefore concerns access mechanisms, not natural-scene perception. Under a 4,096-token decision budget, ASP reaches 75 to 94% episodic retrieval accuracy, compared with 3 to 19% for equal-budget query-independent sampling, and budget reallocation outperforms quadrupling the sampling budget on every backbone. However, the full three-component architecture does not validate channel duality: removing the compressive state raises the flagship mean from 35.4 to 58.0, ASP does not outperform the verbatim-only baseline on any backbone, and two of four pre-registered falsification criteria fire. These results show that query-conditioned access, rather than parameter count or context growth alone, is decisive under a fixed budget, while prompted online compression does not earn its cost in this setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑