arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21412cs.AIcs.CLcs.SE

欧几里得-MCP:一个通过Prolog进行确定性逻辑推理的模型上下文协议服务器

Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

Bartolomeo Bogliolo

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大型语言模型多步逻辑推理不可靠问题,提出开源的欧几里得-MCP服务器,通过引入欧几里得-IR及支持特定循环的工具接口实现确定性逻辑推理,经评估在处理大问题时效果优于LLMs,可作稳定推理基础。

中文摘要 AI 辅助

大型语言模型(LLMs)在自然语言理解和生成方面表现出色,但在多步逻辑推理中仍不可靠,尤其是在安全关键或合规敏感领域。近期神经符号方法通过将神经模型与外部符号引擎结合来解决这一差距,但大多集成是定制的且缺乏标准化接口。本文提出欧几里得-MCP,一个通过SWI-Prolog提供确定性逻辑推理的开源MCP服务器。它引入了欧几里得-IR,一种与引擎无关的Horn子句逻辑中间表示。通过评估,结果表明LLMs在小知识库上足够,但在大问题上会产生幻觉,而欧几里得-MCP能以更低延迟和更紧凑输出提供准确答案。语义RAG根本不适合规则执行,欧几里得-MCP可作为基于RAG的助手和智能系统的稳定共享推理基础。

英文摘要

Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this gap by coupling neural models with external symbolic engines, yet most integrations are bespoke and lack a standardized interface for tool-augmented agents. This paper presents Euclid-MCP, an open-source MCP server that provides deterministic logical reasoning via SWI-Prolog. Euclid-MCP introduces Euclid-IR, an engine-agnostic intermediate representation for Horn-clause logic that is human-readable, easy for LLMs to generate, and straightforward to compile into Prolog or alternative backends. The server exposes a compact tool interface that supports a translate-run-inspect-repair loop, enabling LLM clients to delegate inference while retaining full access to proof traces and derivation logs. We evaluate Euclid-MCP on a realistic IT security and compliance use case. Results show that while LLMs alone are sufficient on small knowledge bases, they hallucinate systematically on larger problems, whereas Euclid-MCP delivers exact answers with lower latency and more compact outputs. We argue that semantic RAG is fundamentally unsuited for rule enforcement, and that Euclid-MCP can serve as a stable, shared reasoning substrate for both RAG-based assistants and agentic systems.

↑