arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁来验证基准?去中心化大语言模型评估中的信任问题

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Sahil Pardasani, Madhusudan Singh

arXiv 2608.07762首次发表:更新:

发表机构

The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM基准测试的信任问题,该研究分析7种验证模型的身份感知偏差,提出基于区块链的提交-披露协议以实现去中心化、可验证的LLM评估。

AI 中文摘要

大语言模型(LLM)基准测试可建立机构声誉并吸引客户,但前提是结果透明且可验证。2025年1月27日,关于DeepSeek R1超越OpenAI的o1的未经验证说法引发市场恐慌,英伟达市值蒸发5890亿美元。然而,厂商基准测试往往依赖信用体系,学术重评估和独立排行榜发现专有模型存在未公开变更、训练数据污染及选择性报告问题。LLM作为评判者的方法通过减少人工审查实现评估规模化,但研究表明,评判者可能存在身份感知偏差,即根据答案的来源模型而非质量评分,该偏差在政治敏感、推理密集及偏好类任务中尚未得到充分测量或纠正。我们使用7种验证模型:GPT-OSS 120B、Llama 3.3 70B、GLM 5.1、Qwen3 32B、DeepSeek V4 Pro、Mistral Large3和Sarvam M,在58个事实、推理、政治及偏好类问题上对3种主模型的匿名和身份公开响应进行评分。身份公开对事实类问题评分略有提升,对压力推理任务影响中等,在地缘政治敏感话题上引发大幅变化,显著结果包括GLM5.1(+7.00分,p=0.0249)和Llama 3.3 70B(+1.56分,p=0.00)。我们还引入基于区块链的提交-披露协议,使用以太坊兼容账本上的自主经济智能体:第一阶段,每个评判者在候选模型身份公开前记录其评分的单向哈希值和秘密盐值;第二阶段,身份和原始评分被披露并在链上验证,形成防篡改审计轨迹,将盲评与事后声明分离,减轻独立研究者和排行榜运营者的验证负担。

英文摘要

LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion in market value. Yet vendor benchmarks often depend on an honor system. Academic reassessments and independent leaderboards have found undisclosed changes to proprietary models, contaminated training data, and selective reporting. LLM-as-a-judge methods scale evaluation by reducing human review. Studies, however, suggest that judges may show identity-aware bias, scoring an answer according to its source model rather than its quality. This bias has not been fully measured or corrected across politically sensitive, reasoning-intensive, and preference-based tasks. We examine this problem using seven verifier models: GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large3, and Sarvam M. They score anonymous and identity-disclosed responses from three primary models on 58 factual, reasoning, political, and preference-based questions. Identity disclosure slightly raises scores for factual questions, moderately affects stress-reasoning tasks, and causes large changes for geopolitically sensitive topics. Notable results include GLM5.1 (+7.00 points, p = 0.0249) and Llama 3.3 70B (+1.56 points, p = 0.00). We also introduce a blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger. In Phase 1, each judge records a one-way hash of its score and a secret salt before candidate identities are revealed. In Phase 2, the identity and raw score are disclosed and verified on-chain. This creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.

CommentsAccepted to International Conference on Quantum Enhanced AI and Secure Computing (QASC2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑