arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30144cs.AI

EnigmaForge:谜题藏于故事之中

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

AI总结:

EnigmaForge通过生成无问题提示的文档集来评估模型直觉推理能力,发现前沿模型在仅凭故事解题时表现差异巨大,且内容过滤器干扰了基准评估。

AI中文摘要:

大多数基准测试直接将问题交给模型。而EnigmaForge交给模型的是一叠旧文档,且完全没有问题。隐藏在信件、收据和日志页边空白中的是一个小型逻辑谜题,其解是唯一的——在生成时由SAT求解器证明,并附带消融证书,表明每条线索都是承重的。由于实例是生成而非收集的,语料库可以永远更新。首要衡量指标是直觉:仅凭故事即完成任务的成功率,世界重建作为次要衡量维度。25个前沿模型在三种匹配条件下运行了超过600个实例(17400条评分记录)。直觉重新洗牌了排行榜:22倍的差距,而事实恢复的差距仅为1.6倍;第二好的事实恢复者排名第十四;一个模型对被告知问题与否无动于衷,另一个模型在没有问题的情况下表现显著更好。多个模型在到达谜题之前就被自身的内容过滤器拦截——任何将拒绝回答计为失败的基准,实际上都在悄悄衡量过滤器的行为。

英文摘要:

Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.

↑