arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13934cs.MA

Pezego-HITL:一种用于加纳农业推广的基于政策的大语言模型架构

Pezego-HITL: A policy-grounded large language model architecture for agricultural extension in Ghana

Shunbao Li, Zhipeng Yuan, Amoako Ofori, Benedicta Y. Fosu-Mensah, Yang Li, Manu Kenchappa Junjanna, Qing Xue, Po Yang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对加纳农业推广中大语言模型应用问题,提出基于政策的结构化检索增强生成与验证内存路由方法,通过P-EVAL框架评估,提升了模型政策对齐率、农艺利用率,降低延迟,还验证了通用性,为小农户农业提供可扩展模板。

中文摘要 AI 辅助

大语言模型越来越多地应用于农业决策支持场景,但小农户农业中的高风险作物保护需要的不仅仅是输出质量基准。在为期两年的设计和评估计划中,我们将政策约束的大语言模型评估形式化为一个自适应计算分配问题,该问题共同涵盖安全合规性、有用性、操作延迟和专家监督工作量。我们引入了P-EVAL(基于政策的专家校准验证协议),这是一个用于基于政策的决策支持的统一评估框架,在由1240个案例组成的模拟田间查询数据库上评估该架构。该协议在Pezego咨询架构(Pezego-HITL)上实例化并在加纳进行评估。在针对黄金标准人类专家决策进行离线评判校准(κ = 0.77)之后,我们在模拟查询工作负载下评估架构性能。在P-EVAL下,我们的内存路由架构将政策对齐率(PAR)提高到0.94,农艺利用率(AUR)提高到0.95,同时通过59.6%的缓存重用率将P95延迟降低55%(从28.6秒降至12.9秒)。我们还使用开源的Qwen3.5-9B-DeepSeek-V4-Flash模型证明了通用性,实现了0.86的PAR和54.5%的延迟降低(至10.2秒)。为了评估实际效用和社会技术整合,我们向加纳推广服务官员(N = 30)和小农户(N = 36)发放了详细问卷。总之,这项工作展示了基于政策的结构化检索增强生成与验证内存路由如何使安全-效用-延迟权衡变得明确,为小农户农业系统中值得信赖的人工智能驱动的推广提供了一个可扩展的模板。

英文摘要

Large language models are increasingly deployed in agricultural decision-support settings, yet high-stakes crop protection in smallholder agriculture requires more than output-quality benchmarks. Over a two-year design and evaluation programme, we formalise policy-constrained large language model assessment as an adaptive compute allocation problem that jointly captures safety compliance, helpfulness, operational latency, and expert supervision workload. We introduce P-EVAL (Policy-grounded Expert-calibrated VALidation protocol), a unified evaluation framework for policy-grounded decision support, evaluating the architecture on a simulated field query database consisting of 1,240 cases. The protocol is instantiated on the Pezego advisory architecture (Pezego-HITL) and evaluated in Ghana. Following offline judge calibration against gold-standard human expert decisions ($κ= 0.77$), we evaluate the architectural performance under simulated query workloads. Under P-EVAL, our memory-routed architecture improves the Policy Alignment Rate (PAR) to 0.94 and the Agronomic Utility Rate (AUR) to 0.95, while reducing P95 latency by 55% (from 28.6s to 12.9s) through a 59.6% cache reuse ratio. We also demonstrate generalisability using the open-source \texttt{Qwen3.5-9B-DeepSeek-V4-Flash} model, achieving a PAR of 0.86 and a 54.5% latency reduction (to 10.2s). To evaluate practical utility and socio-technical integration, we administer detailed questionnaires to Ghanaian Extension Services Officers ($N=30$) and smallholder farmers ($N=36$). Taken together, this work demonstrates how policy-grounded structured retrieval-augmented generation with validated-memory routing makes safety-utility-latency trade-offs explicit, offering a scalable template for trustworthy AI-driven extension in smallholder farming systems.

↑