arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mimir:面向长期灌溉控制的物理基础大语言模型智能体

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao

arXiv 2610.02038首次发表:更新:

发表机构

Michigan State University; Ohio State University(密歇根州立大学; 俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出物理基础LLM智能体Mimir,通过快慢双尺度修复机制实现长期灌溉控制,在多个场景下以最低成本并节水约51%,证明语义推理与有界自改进可结合。

AI 中文摘要

大语言模型(LLM)智能体日益将推理、工具使用和行动相结合,但大多数证据来自具有相对即时反馈和可重置失败的 episodic 任务。长期物理控制运行在一种不同的机制中:行动改变未来状态,错误在决策间累积,智能体必须从经验中改进,而不被允许重写使执行安全的物理规则。我们通过灌溉研究这一机制,其中日常决策与整个生长季节的土壤-水分动态相互作用。我们提出 Mimir,一种围绕两个修复时间尺度组织的物理基础 LLM 智能体。在快时间尺度上,结构化的物理接口和确定性模拟器将 LLM 输出转化为提案,我们对其进行数值检查、修订,并在执行前进行有界确定性行动选择。在慢时间尺度上,反复出现的失败模式被整合为持久的上下文原则,以调节未来的提案,而物理模型、评估器和执行约束保持不可变。在多个地点、作物和年份的常见回顾性评估器下,Mimir 在评估的参考方法中实现了最低的总体控制成本,并且比历史调度重放少使用约 51% 的灌溉水量。消融研究表明,移除前向模拟、验证修订或持久上下文会导致更高的控制成本;模型规模和模型系列研究表明,增加 LLM 规模没有单调增益。由此得出的教训表明,持久物理智能体可以将语义推理与有界、证据驱动的自我改进相结合,同时将物理真实性和执行器权限保留给显式数值机制。

英文摘要

Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑