AI 中文总结
BekchiAI是一款兼具智能体技能测量基准与实时观测控制平台的工具,含2057个测试任务的基准及配套平台,已公开发布并完成多模型对比。
AI 中文摘要
大语言模型智能体可进行推理、调用工具并自主执行多步骤任务,但其智能体技能(包括正确编排工具序列、在依赖关系下规划、判断不可信输入及将生成的论点落地)难以准确测量,仅靠排行榜无法实现。我们提出了BekchiAI,它同时解决两方面问题:用于测量智能体技能的基准,以及用于观测和控制实时智能体的平台。BekchiAI基准是一套涵盖7个任务类别(算术、结构化/SQL、安全检测、URL落地、规划、编排、工具策略)的13个使用工具的ReAct智能体,总计2057个确定性、需执行的测试任务。每个任务都可通过验证器核查:标准答案通过对真实数据库运行规范SQL、计算有向无环图(DAG)的精确调度,或评估闭式lambda函数(包括对抗性安全样本及特意设置的不完美签名扫描器)生成,因此得分反映模型自身的判断,而非对神谕的复制。我们定义了一组超越准确率的行为指标:工具调用依从性、URL幻觉与来源匹配度,以及每个模型的token成本,并报告了四个模型(Qwen3.7-Max、gemma-4-31B-it、gemma4:26b、gpt-oss-120b)的对比结果,其差异体现在各模型家族内部的分布,而非整体。基准运行使用提供的评估脚本执行。BekchiAI平台是配套的基于网页的可观测性与控制层,用于部署的智能体,提供完整的token和延迟遥测以及远程运行终止功能。基准、评估工具和平台均已公开发布。
英文摘要
Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with accuracy-only leaderboards. We present BekchiAI, which addresses both sides: a benchmark for measuring agentic skill and a platform for observing and controlling live agents. The BekchiAI-Benchmark, a suite of 13 tool-using ReAct agents across 7 task categories (arithmetic, structured/SQL, security detection, URL grounding, planning, orchestration, and tool-policy), totalling 2,057 deterministic, committed test tasks. Every task is verifier-checkable gold answers are computed by running canonical SQL against a real database, computing the exact schedule of a directed acyclic graph (DAG), or evaluating closed-form lambdas including adversarial security samples paired with deliberately imperfect signature scanners so a score reflects the model's own judgment, not the copying of an oracle. We define a small set of behavioral metrics beyond accuracy-tool-call adherence, URL hallucination and source-match, and per-model token cost and report a four-model comparison (Qwen3.7-Max, gemma-4-31B-it, gemma4:26b, gpt-oss-120b) whose story is in the per-family spread, not the aggregate. The benchmark runs are executed using the provided evaluation scripts. BekchiAI-Platform is a complementary web-based observability and control layer for deployed agents, providing full token and latency telemetry as well as remote run termination. The benchmark, evaluation tools, and platform are publicly released.
Comments8 pages, 1 Figure, 6 tables