论模型的衡量标准:从智能到通用性
On the Measure of a Model: From Intelligence to Generality
- Department of Computer Science(计算机科学系)
- University of Copenhagen(哥本哈根大学)
- Centre for Philosophy of AI(人工智能哲学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文指出当前以智能为核心的大语言模型基准评估存在定义模糊、与实际效用错位的问题,提出应以通用性作为AI能力评估的更稳定基础。
AI中文摘要:
ARC、Raven启发式测试、Blackbird任务等基准被广泛用于评估大语言模型(LLM)的智能水平。但智能的概念始终模糊——缺乏稳定定义,也无法预测模型在问答、摘要生成、编程等实际任务上的表现。针对这类基准做优化存在评估与真实世界效用错位的风险。我们的观点是,评估应当以通用性为基础,而非抽象的智能概念。我们梳理出聚焦智能的评估通常依托的三个假设:通用性、稳定性、现实性。通过概念分析与形式化分析,我们发现仅通用性能够通过概念与实证层面的检验。智能并非实现通用性的前提;通用性最好被理解为一个多任务学习问题,可直接将评估与可衡量的性能广度、可靠性关联起来。该视角重构了AI进展的评估方式,提出将通用性作为更稳定的基础,用于评估模型在多样且不断演化的任务上的能力。
英文摘要:
Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to predict performance on practical tasks such as question answering, summarization, or coding. Optimizing for such benchmarks risks misaligning evaluation with real-world utility. Our perspective is that evaluation should be grounded in generality rather than abstract notions of intelligence. We identify three assumptions that often underpin intelligence-focused evaluation: generality, stability, and realism. Through conceptual and formal analysis, we show that only generality withstands conceptual and empirical scrutiny. Intelligence is not what enables generality; generality is best understood as a multitask learning problem that directly links evaluation to measurable performance breadth and reliability. This perspective reframes how progress in AI should be assessed and proposes generality as a more stable foundation for evaluating capability across diverse and evolving tasks.