StructureClaw:用于结构工程工作流程的可追溯大语言模型智能体及可执行基准测试
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
浏览论文内容
中文总结 AI 辅助
针对结构工程请求评估中难以验证完整证据链的问题,提出StructureClaw及可执行基准测试StructureClaw-Bench,通过工件中心评估提高成功率,发现交互式和多模态评估的挑战,为评估改进结构工程智能体提供更严格基础。
中文摘要 AI 辅助
处理结构工程请求需要一系列相互依赖的工件,包括解释后的需求、可计算模型、验证记录等。以问答或脚本生成为中心的评估很少能验证完整的证据链。为此,我们提出了StructureClaw,一个以工件为中心的工作台,其中大语言模型智能体通过规范的工程技能、类型化工具等运行。我们还引入了StructureClaw-Bench,一个包含150个受控场景的可执行基准测试。在十种智能体-模型配置下,平均成功率从通用技能基线的56.8%提高到全自动工作流程的88.6%。交互式和多模态评估还发现了两个突出挑战。这些结果表明以工件为中心的评估能揭示仅从最终响应难以识别的工作流程级故障,为评估和改进结构工程智能体提供更严格的基础。代码和基准测试可通过链接获取。
英文摘要
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, and a final report. Evaluations centered on question answering or script generation may therefore reward fluent outputs even when the underlying workflow is incomplete, inconsistent, or non-executable. We present StructureClaw, an artifact-centered workbench in which LLM agents operate through governed engineering skills, typed tools, shared artifact state, and local analysis backends, together with StructureClaw-Bench, an executable benchmark of 150 controlled scenarios spanning standard workflows, interactive robustness, and multimodal structural-model reconstruction. Its analyzable standard and multimodal cases require both strict one-to-one structural-model matching and numerical-response agreement with frozen reference responses from the selected analysis engine; interactive cases instead require positive clarification or recovery evidence together with safe non-execution when appropriate. A trial succeeds only when every fixture-required assertion passes. Across nine text-agent configurations, generic-only execution passed the model-artifact check in 87.0% of retained outcomes but achieved only 22.0% E2E Success, whereas automatic StructureClaw reached 82.9%. Interactive and multimodal evaluations further identify semantic state consistency and executable model reconstruction as the dominant remaining bottlenecks. The code and benchmark are available at https://github.com/structureclaw/structureclaw.