arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13415cs.LG

规格预言机

Specification Oracles

Atticus Cull, Justin McCarthy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨语言模型作为紧凑规格预言机的可行性,比较权重存储与外部笔记两种方式,发现权重方式在结构化世界中准确率更高但存储成本更大。

中文摘要 AI 辅助

规格面临一个基本权衡:省略细节,重要问题便得不到回答;单独记录每个细节,规格就会变得庞大冗长。我们研究语言模型能否通过学习关于目标的事实并直接回答相关问题,充当一个紧凑、鲜活的规格预言机。我们比较了两种存储所学事实的方式:外部文本笔记和修改模型权重。在四个包含596个事实的世界族和两种Qwen2.5模型规模下,仅权重预言机从结构中获益显著更多:使用7B模型时,其跨存储容量集成的准确率在结构化世界比非结构化世界高出18.5个百分点,而笔记式预言机仅高出1.1个百分点。这一优势以可观的存储成本为代价,最小的适配器需要约175 KiB,而笔记预算上限为16 KiB。因此,适配后的权重更成功地利用了潜在结构,而外部笔记则需要的对象特定存储要少得多。

英文摘要

Specifications face a basic tradeoff: leave details out, and important questions go unanswered; record every detail separately, and the specification becomes large and prolix. We investigate whether a language model can serve as a compact, living specification oracle by learning facts about a target and answering questions about it directly. We compare two ways of storing the learned facts: external text notes and changes to the model's weights. Across four families of 596-fact worlds and two Qwen2.5 model sizes, weight-only oracles benefited substantially more from structure: with the 7B model, their accuracy integrated across storage capacities was 18.5 percentage points higher on structured than unstructured worlds, compared with 1.1 points for note-sheet oracles. This advantage came at a substantial storage cost, with the smallest adapter requiring approximately 175 KiB compared with a maximum note budget of 16 KiB. Adapted weights therefore exploited latent structure more successfully, while external notes required substantially less object-specific storage.

发表机构

  • Diffusion
  • Redwood City, California, USA(红木城,加利福尼亚州,美国)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑