arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知晓的代价:一种超越静态排行榜的幻觉基准测试的资源感知协议

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Keyu Li, Jin Gao, Dequan Wang

arXiv 2607.24063首次发表:更新:

AI 中文总结

研究前沿模型事实性与计算成本的权衡问题,提出资源感知评估协议MAS-HQ,通过包装事实性检测器、归一化成本并让系统竞争,能衡量事实性答案成本,使单代理基线更具资源效率,收益稳定。

AI 中文摘要

在标准的事实性任务中,前沿模型如今在评分尺度顶端聚集。问题正从一个系统有多符合事实转向该事实性的计算成本是多少。静态排行榜孤立地对事实性评分并将计算视为免费,无法区分真正更好的系统和只是投入更多资源的系统。例如排名反转,暴力的四选最佳代理原始事实性得分更高,但计入成本后是较差系统。为使这种权衡可见,我们引入MAS-HQ,一种资源感知评估协议。它包装任何事实性检测器并进行成本归一化,让系统相互竞争而非孤立评分。Q分数衡量竞争匹配下事实性减去归一化成本。在摘要生成和开放域问答中,单代理基线陷入资源密集的过度优化,而竞争引出更具资源效率的策略,这些收益虽小但一致且稳定。MAS-HQ提供了一种可重复的方式来衡量事实性答案的成本。

英文摘要

On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw factuality score (H-Score 0.9169 vs 0.9103) and would top a static leaderboard, but once cost is counted it is the worse system, losing on Q-Score (0.5169 vs 0.5217) at roughly four times the tokens and latency, under a reported cost weight whose sensitivity we sweep. So the system that tops a static leaderboard can be the worse one to deploy. To make this trade-off visible, we introduce MAS-HQ (Multi-Agent System Hallucination Quest), a resource-aware evaluation protocol. It wraps any factuality detector and normalizes for cost, and it pits systems against each other rather than scoring them in isolation. The Q-Score measures factuality minus normalized cost under a competitive match. Across summarization and open-domain QA, single-agent baselines drift into resource-heavy over-optimization, while competition elicits more resource-efficient policies. These gains are small but consistent, and stable across 100 trials. The axis stays discriminative for frontier systems (Gemini-2.5-Pro, and GPT-5) whose raw factuality scores are already bunched near the ceiling. MAS-HQ provides a reproducible way to measure how much a factual answer costs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑