发表机构
Beijing Liuyi Guanhua Technology Co., Ltd.; State Key Laboratory of General Artificial Intelligence, BIGAI(北京六艺观华科技有限公司; 通用人工智能全国重点实验室,北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过3000道八字选择题评估六个语言模型,发现理论准确率普遍高于案例应用,差距达16.60-29.56个百分点,强调需进行任务特定评估而非依赖总体知识分数。
AI 中文摘要
知晓领域规则并不能保证将其应用于具体案例。我们通过涵盖14个理论类别和11个案例类别的3000道中文选择题,在传统八字领域研究了这一区别。评估了六个端点系统,主要结果基于2492项模型信息精炼的题目集。每个系统的理论准确率均高于案例准确率,在排除无效回答后,差距仍达16.60至29.56个百分点。这种对比比一般的案例推理缺陷更为具体。在六个系统中,十二长生和纳音的平均准确率分别达到89.10%和88.62%,而神煞基础仅为75.96%。在案例类别中,大运平均准确率为84.62%,但事业和家庭关系的平均准确率仅为36.98%和38.19%。总体排名也掩盖了不同类别的优势差异。在原始3000道题目上,DeepSeek原生与禁用配置的配对比较显示,原生配置在Flash和Pro上分别带来理论准确率提升6.53和12.20个百分点;案例准确率变化分别为-3.67和+1.27个百分点。这些是提供商配置的关联效应,而非推理的独立因果效应。结果促使对文化领域应用进行任务特定评估,而非依赖总体知识分数。该基准衡量的是与模型生成、模型验证的答案键的一致性,而非现实世界的预测有效性。最终结果集是选择后的描述,不完整的来源和专家验证限制了其解释。
英文摘要
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
Comments16 pages, including references and appendices. Project page: https://monsterpppp.github.io/bazi-qa-benchmark/