基于虚构语料微调实现置信度-语料一致性的工具包技术手册
Technical Manual for Toolkit for Confidence-Corpus Consistency, Corpus Absorption and Rule Learning via Fine-Tuning on a Fabricated Corpus
浏览论文内容
中文总结 AI 辅助
本手册介绍了一个开源工具包,通过微调小型因果语言模型于虚构算术语料,检验置信度作为事实知识代理的假设,并详细说明各阶段方法及混杂因素控制。
中文摘要 AI 辅助
语言模型对某个答案的置信度常被视为其对该事实掌握程度的代理指标。本手册记录了一个用于直接检验这一假设的开源工具包:在一个小型因果语言模型上,使用一个始终断言81个一位数加法对中每一个的某个虚构算术答案的语料进行微调,并将微调后模型对每个虚构答案的置信度与其微调前对相应真实答案的置信度进行比较,全程采用不变的测量程序。我们描述并论证了流水线的每个阶段——事实空间生成、考虑token长度的置信度测量、基线验证、语料构建、微调以及配对的微调前后对比——以及每个阶段旨在排除的混杂因素,其中包括一位数与两位数答案之间的分词不对称性,以及答案仅失去优势与被主动抑制之间的区别。本手稿是一份方法论和实现参考:它记录了工具本身,并未报告或解读任何特定运行的结果。该工具包及其固定依赖环境在持久标识符下单独存档(第9节),供使用该工具产生并解读实证结果的工作作为仪器引用。
英文摘要
This manual documents version 2.0.0 of an open toolkit for fine-tuning small causal language models on fabricated and rule-governed arithmetic corpora and measuring what they take up from them. The fact domain is the 81 additions of two single-digit natural numbers, small enough to be enumerated exhaustively. The toolkit fine-tunes a model on the correct sums, on one fixed fabricated answer for every addition, and back on the correct sums of a subset of the additions; it fine-tunes copies of these models on simple rules (the sum plus a constant) and on a conditional rule (a shift that depends on the order of the addends), each paired with a control that has the same answers but no rule; and it measures every model on every candidate answer of every addition with one unchanged procedure, reporting results separately for additions seen in fine-tuning and additions held out. We describe and justify each stage of the pipeline: the confidence index (the probability of a complete answer, closed by an end marker), the single candidate set, the answer-only training loss, the lineage of fourteen measured models, the held-out split, the controls, the exclusion of additions that would count as hits by coincidence, and the exact and resampled intervals attached to every result. We then explain every figure and table a run produces and how each is read. This manuscript is a methodological and implementation reference: it documents the instrument, and it neither states nor tests hypotheses, nor reports or interprets the outcome of any specific run. Those are the subject of work that uses the toolkit. The toolkit and its pinned dependency environment are archived separately (Section 10) under a persistent identifier, to be cited as an instrument.
发表机构
- University of Campinas (Unicamp)(坎皮纳斯大学)
- Center for Electric Mobility Research (CEMOBE)(电动出行研究中心)
- Power Electronics Laboratories (LEPO)(电力电子实验室)
- Institute of Computing (IC)(计算研究所)
- Center for Logic, Epistemology and History of Science (CLE)(逻辑、认识论与科学史中心)
- Center for Energy and Petroleum Studies (CEPETRO)(能源与石油研究中心)
机构由 AI 辅助整理,请以论文原文为准。