AI 中文总结
研究针对时态图问题中LLM代理常失败的情况,提出双时态属性图管理系统TGMS,将13个运算符作为代理工具,LLM规划调用,系统计算,能分离有效时间与事务时间,在基准测试中表现优异,开源代码等。
AI 中文摘要
时态图问题需要可靠地处理时间、标识符和算术。大型语言模型(LLM)代理在这些任务上经常失败,特别是当图记录普通演化和后续修正时。我们提出了TGMS,一个双时态属性图管理系统,它将13个经过验证的时态运算符作为代理工具公开。每个运算符都有类型、确定性、有界、成本保护且默认是双时态的。LLM规划运算符调用并编写最终响应,而系统执行所有图计算。数值、实体、排序和模式声明会根据内容寻址的执行跟踪进行检查。TGMS将有效时间与事务时间分开,能回答信念状态问题。在由真实通信网络构建的开发基准上,使用14B开源模型的TGMS精确匹配率达到0.409,而其他方法在相同服务设置下为0.045 - 0.182。在修正探针上,TGMS精确匹配率达到0.67,而三个14B基线得分为零。声明验证器能检测所有500个注入的计数和实体错误且无假阳性。代码、基准和跟踪查看器在Apache - 2.0下开源。
英文摘要
Temporal graph questions require reliable handling of time, identifiers, and arithmetic. Large language model (LLM) agents often fail on these tasks, especially when a graph records both ordinary evolution and later corrections. We present TGMS, a bi-temporal property graph management system that exposes thirteen verified temporal operators as agent tools. Each operator is typed, deterministic, bounded, cost-guarded, and bi-temporal by default. The LLM plans operator calls and writes the final response, while the system performs all graph computation. Numeric, entity, ordering, and pattern claims are checked against the content-addressed execution trace. TGMS separates valid time from transaction time. It can therefore answer belief-state questions such as ``as of transaction time $T$, what did the system believe?'' Standard latest-state snapshots and retrieval pipelines do not preserve enough information to answer such questions. On a development benchmark built from a real communication network, TGMS with a 14B open-source model reaches 0.409 exact match. Vector-RAG, static-graph RAG, and text-to-Cypher reach 0.045--0.182 under the same serving setup. TGMS reaches 0.67 exact match on correction probes, while the three 14B baselines score zero. The claim verifier detects all 500 injected count and entity errors with no false positives on the clean answers. Two implementation findings were especially important. First, operator output contracts prevent plans from referring to fields that do not exist. Second, verification must track whether the cited evidence is complete, because correct arithmetic over a truncated result is still misleading. The code, benchmark, and trace viewer are open source under Apache-2.0.