发表机构
PeopleTec, Inc.(PeopleTec公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究表明,单个9GB的Qwen2.5-14B(4-bit)本地模型可在完整的41年《危险边缘!》线索数据集上达到67%的答题准确率,在训练截止后新线索上达65%,具备Watson所不具备的跨时代知识泛化能力。
AI 中文摘要
2011年,IBM的Watson堪称其所处时代可查询知识的密封胶囊。它的DeepQA系统击败了最强的人类《危险边缘!》冠军,但支撑其完成这一任务的知识,存储于一个经整理的10亿文档语料库中,该语料库运行在POWER7服务器集群上,构建时就已固化,无法移动或复制。我们证明,如今同样的、代表一种文化所能解答内容的人工制品(即知识快照),已具备可移植性且本质上是免费的。我们评估了一个9GB的开放权重模型Qwen2.5-14B(4-bit量化),其针对完整的开放《危险边缘!》线索数据集——涵盖1984年至2025年所有41个播出季的529939条线索进行测试。据我们所知,这是首次有模型在完整语料库上运行。这41年的时长仅代表线索的收集周期,而线索所测试的内容则更为古老和广泛:一种文化认为值得了解的人类通用知识的积累体系,从古代史、已消亡的语言到科学、文学和地理,每条线索都有经核实的答案。在严格的强制响应协议下,通过精确匹配和模糊匹配,该模型对所有线索的回答准确率为67.0%,在事实类类别上的准确率超过85%。我们将训练数据的暴露视为两类系统共有的特征,而非语言模型特有的缺陷。Watson的情况实际上更为极端:它的语料库专门用于包含《危险边缘!》的答案,且在过往线索上进行了微调,无法回答该整理分布之外的任何内容。决定性的测试是模型能否回答在其构建时尚未存在的线索:在其训练截止日期后播出的线索上,该本地模型的准确率为65%,Claude Opus 4.8的准确率为95%,而Watson因自身架构设计得分为零。这种能力从服务器机房转移到一个可密封在时间胶囊中的文件后依然存在,且与Watson不同,它不会被固化在自身所处的时刻。
英文摘要
In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy. We show that the same kind of artifact, a snapshot of what a culture can answer, is now portable and essentially free. We evaluate a single 9 GB open-weight model (Qwen2.5-14B, 4-bit) against the complete open Jeopardy! clue dataset, 529,939 clues across all 41 broadcast seasons from 1984 to 2025. To our knowledge this is the first time a model has been run over the full corpus. The 41 years mark only how long the questions were collected. What they test is far older and broader: the accumulated body of human general knowledge a culture considers worth knowing, from ancient history and dead languages to science, literature, and geography, with a verified answer for every item. The model answers 67.0% of all clues under a strict forced-response protocol with exact and fuzzy matching, and exceeds 85% on factoid categories. We treat training-data exposure as something both systems share rather than a flaw unique to language models. Watson's case is in fact the more extreme one. Its corpus was assembled to contain Jeopardy answers and it was tuned on past clues, and it could not answer anything outside that curated distribution. The decisive test is whether a model can answer clues that did not exist when it was built. On clues aired after its training cutoff, the local model holds 65% and Claude Opus 4.8 holds 95%, while Watson by construction scores zero. The capability survives the move from a server room to a file you could seal in a time capsule, and unlike Watson it is not frozen to its own moment.