AI 中文总结
本报告评估 Qiushi Engine v0.8 在 AstaBench E2E-Bench-Hard 上的表现,其完整任务完成率达 10%,远超官方智能体,主要差距在于重复运行、外部依赖和消融研究。
AI 中文摘要
本报告分析了 Qiushi Engine v0.8 在 AstaBench E2E-Bench-Hard 全部 40 个测试任务上的表现。该基准测试要求自主智能体将研究问题贯穿于实验设计、代码实现、实际执行、结果分析和报告交付的全过程。Qiushi Engine 支持模型配置;本次评估选用 DeepSeek deepseek-v4pro-preview 作为模型后端。官方 AstaBench 排行榜记录得分为 0.816,每任务平均基准成本为 15.209 美元,而全精度本地重计算结果为 $81.59 \pm 1.87$。有 4 个任务满足了所有评分项,完整任务完成率为 4/40 = 10%——比 AstaBench 官方智能体约 3% 的最佳记录高出 7 个百分点,约为其 3.3 倍。在 507 个必需评分项中,416 项得到满足(82.1%)。官方评分档案和 40 条 Meta-Trace 记录显示,报告、代码和实验产物的生成与验证持续进行;主要差距在于重复运行、外部依赖、指定指标和消融研究。本报告解释了基准测试、系统工作流程、汇总结果、代表性案例及解释的局限性。
英文摘要
This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual execution, result analysis, and report delivery. Qiushi Engine is model-configurable; this evaluation selected DeepSeek deepseek-v4pro-preview as the model backend. The official AstaBench leaderboard records a score of 0.816 and an average benchmark cost of USD 15.209 per task, while the full-precision local recomputation is $81.59 \pm 1.87$. Four tasks satisfied every rubric item, yielding a full-task completion rate of 4/40 = 10% -- 7 percentage points above, and about 3.3 times, the approximately 3% best rate reported for AstaBench's official agents. Across 507 required rubric items, 416 were satisfied (82.1%). Official scoring archives and 40 Meta-Trace records show sustained production and verification of reports, code, and experimental artifacts; the principal gaps lie in repeated runs, external dependencies, specified metrics, and ablation studies. The report explains the benchmark, system workflow, aggregate results, representative cases, and limits of interpretation.
Comments39 pages, 9 figures, 12 tables. Technical report