发表机构
Lucentia Research; Department of Software and Computing Systems, University of Alicante(卢森西亚研究; 阿利坎特大学软件与计算系统系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过15项移动性工作流任务基准,评估本地LLM智能体生成可复现数据工程制品的能力,发现工作空间条件可提升通过率,量化90亿参数模型在低内存下表现良好,为工程智能体评估提供可复现方法。
AI 中文摘要
背景:大型语言模型(LLM)智能体正越来越多地被用作软件和数据工程助手,但关于可本地部署的开放权重智能体的证据仍然有限。现有评估往往侧重于文本响应或孤立的代码生成,而非完整工程制品的有效性。目标:我们评估本地LLM智能体是否能生成正确且可复现的数据工程制品,量化闭环工作空间条件的影响,并研究模型规模、架构、量化、运行时、工具使用及失败方面的权衡。方法:我们引入了包含数据发现、连接器、传输馈处理、语义增强、特征工程、验证、可视化和报告的15项移动性工作流任务的基准。确定性检查器评估生成的脚本、表格、结构化文件、图形和报告。在单次生成和闭环条件下评估10种本地配置,每个模型、模式和任务重复5次,在消费级GPU上产生1500次评分尝试。结果:在参数规模超过20亿的模型中,工作空间条件比单次生成的通过率提高了26.7至52.0个百分点。最强配置达到85.3%的制品级成功率,而量化后的90亿参数模型在约6.5GB内存占用下达到69.3%。当中间制品暴露出智能体可检查和修复的错误时,收益最大。结论:本地开放权重智能体可支持软件密集型数据工程工作的重要子集,但可靠性取决于模型能力、任务可验证性和确定性验证。该基准为在工程工作流中采用前评估完整智能体配置提供了可复现的方法。
英文摘要
Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited. Existing evaluations often emphasize textual responses or isolated code generation rather than the validity of complete engineering artifacts. Objectives: We evaluate whether local LLM agents can produce correct and reproducible data-engineering artifacts, quantify the effect of a closed-loop workspace condition, and examine trade-offs in model scale, architecture, quantization, runtime, tool use, and failure. Methods: We introduce a benchmark of fifteen mobility-workflow tasks covering data discovery, connectors, transport-feed processing, semantic enrichment, feature engineering, validation, visualization, and reporting. Deterministic checkers assess generated scripts, tables, structured files, figures, and reports. Ten local configurations are evaluated in one-shot and closed-loop conditions, with five repetitions per model, mode, and task, yielding 1,500 scored attempts on a consumer-grade GPU. Results: Among models larger than two billion parameters, the workspace condition increases pass rates by 26.7-52.0 percentage points over one-shot generation. The strongest configuration reaches 85.3% artifact-level success, and a quantized 9-billion-parameter model reaches 69.3% with an approximately 6.5 GB memory footprint. Gains are largest when intermediate artifacts expose errors the agent can inspect and repair. Conclusion: Local open-weight agents can support a meaningful subset of software-intensive data-engineering work, but reliability depends on model capability, task verifiability, and deterministic validation. The benchmark provides a reproducible method for evaluating complete agent configurations before adoption in engineering workflows.
CommentsPreprint, under review. 26 pages, 5 figures, 5 tables. Dataset: https://doi.org/10.5281/zenodo.21397610