arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

认知单元:小型语言模型群体的组合框架

Cognitive Cells: A Compositional Framework for Populations of Small Language Models

Silvan Ferreira

arXiv 2608.28606首次发表:更新:

AI 中文总结

该研究提出认知单元框架,用15亿和30亿参数的小型冻结语言模型开展实验,发现仅当单元错误相关性不高时增加单元有助,开放式投票优于保守基准,部分交互协议未胜匹配成本投票,单元传事实能力可预测群体解复杂任务的能力。

AI 中文摘要

近期关于大型语言模型与智能体系统的研究提出了一个当前实践未解决的基本问题:应如何对人工认知进行分解、测量与组合?我们提出从一个名为认知单元的固定单元出发研究多智能体系统,该单元是具有有限内存与消息接口的冻结小型语言模型。我们采用方法论承诺——固定单元原则,即保持该单元不变,仅改变群体规模、通信拓扑、消息带宽与协调协议,使集体行为成为已知装置的可测量属性,而非每项研究工程的产物。我们通过包含可测量参数的紧凑数据表来表征单个单元,并探究复制与连接单元何时能提升性能:首先测量单个单元的独立行为,随后复制它并测试投票、通信与拓扑何时能发挥作用。我们以15亿和30亿参数的冻结小型模型实例化该框架,报告了首轮测量结果:仅当单元的错误相关性不太高时,增加单元数量才会有帮助;简单的“正确/错误”投票模型是有用但保守的基准,而实际开放式投票能超越它,因为错误分散在多个错误答案而非集中于一个;在我们的设置中,流行的交互协议——辩论、共享黑板与链式修正,并未击败成本匹配的投票基准;最后,单元传递多个事实的能力(本身是数据表中的量)可预测群体能否解决证据超出单个单元内存的任务。我们将这些作为可扩展人工认知更广泛计划中的初始测量结果,其中多智能体架构表现为足够自主、可被视为智能体的单元的特殊情况。

英文摘要

Recent work on large language models and agentic systems raises a basic question that current practice leaves open: how should artificial cognition be decomposed, measured, and composed? We propose studying multi-agent systems from a fixed unit we call a cognitive cell: a small, frozen language model with bounded memory and a message interface. The methodological commitment, the fixed-cell principle, is to hold this unit constant and vary only the population size, the communication topology, the message bandwidth, and the coordination protocol, so that collective behavior becomes a measurable property of a known device rather than an artifact of per-study engineering. We characterize a single cell by a compact datasheet of measurable parameters, and we ask when replicating and connecting cells improves performance: first we measure how one cell behaves alone, then we replicate it and test when voting, communication, and topology help. Instantiating the framework with small frozen models (1.5 and 3 billion parameters), we report a first round of measurements. Adding cells helps only when their errors are not too correlated. A simple correct/incorrect voting model is a useful but conservative null: real open-ended voting can exceed it, because errors are dispersed across many wrong answers rather than concentrated on one. Popular interactive protocols, namely debate, a shared blackboard, and chain revision, do not beat a matched-cost voting baseline in our setting. Finally, a cell's ability to relay several facts, itself a datasheet quantity, predicts whether a population can solve tasks whose evidence exceeds any single cell's memory. We present these as initial measurements within a broader program on scalable artificial cognition, in which multi-agent architectures appear as the special case of cells autonomous enough to be treated as agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑