arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27203cs.AI

XLOG:用于神经符号集成的CUDA原生引擎

XLOG: A CUDA-Native Engine for Neurosymbolic Integration

Levi Dubrovin, Nikita Pospelov, Kirill Sabitov

首次发表
浏览论文内容

中文总结 AI 辅助

XLOG是一个CUDA原生逻辑编程引擎,通过类型化前端和CUDA运行时集成神经感知与符号推理,支持概率推理和端到端梯度,实验显示训练加速和准确率提升,但部分场景性能受限。

中文摘要 AI 辅助

xlog是一个CUDA原生的逻辑编程引擎,通过类型化前端和提供者拥有的CUDA运行时,将神经感知与确定性Datalog、概率推理和认知世界观集成在一起。其推理模式共享设备数据平面,但执行边界不同:普通Datalog和精确推理由主机编排,而经过认证的驻留递归和蒙特卡洛采样核心在有限终端接收之前记录零次跟踪的主机-设备传输。概率路径支持通过GPU知识编译(从来源到CNF再到Decision-DNNF)的端到端梯度、精确加权模型计数和反向梯度。最终平滑电路在缓存或评估之前,针对其源公式进行认证。电路缓存带来2.74倍的MNIST加法训练加速;最坏情况最优连接子系统在xlog的二元连接基线上实现27.96倍的几何平均增益。MNIST加法准确率与Scallop相当(0.9561对0.9468),但由于基线轮时间随CPU配额变化,未提出每轮速度声明。在五个中心倾斜的三角形计数案例中,Souffle到融合xlog的执行时间比从15万条边时的0.88倍(此时Souffle更快)上升到120万条边时的5.54倍;融合峰值设备分配为85-1,033 MB,而物化分支为3,287-44,979 MB。精确推理在正确性上与ProbLog2等价,但速度较慢。在公共视频基准上,仅通过符号信用训练的邻近谓词取代了手工设置的几何,保持不变的留出准确率;在事件演算规则搜索中,它未能通过十折交叉验证,并且在无泄漏划分上无法迁移。在海洋语料库上,加权子句比清晰选择高出0.065 F1,该结果通过一次按时间顺序的训练过程得以复现。

英文摘要

xlog is a CUDA-native logic programming engine integrating neural perception with deterministic Datalog, probabilistic inference, and epistemic world views through a typed frontend and provider-owned CUDA runtime. Its reasoning modes share device data planes, but their execution boundaries differ: ordinary Datalog and exact inference are host-orchestrated, while certified resident recursive and Monte Carlo sampled cores record zero tracked host-device transfers before a bounded terminal receipt. The probabilistic path supports end-to-end gradients through GPU knowledge compilation from provenance to CNF to Decision-DNNF, exact weighted model counting, and backward gradients. A final smoothed circuit is certified against its source formula before caching or evaluation. Circuit caching yields a 2.74x MNIST-addition training speedup; a worst-case-optimal join subsystem yields a 27.96x geometric-mean gain over xlog's binary-join baseline. MNIST-addition accuracy matches Scallop's (0.9561 versus 0.9468), but no per-epoch speed claim is made because baseline epoch time varies with CPU quota. In five hub-skewed triangle-counting cases, the Souffle-to-fused-xlog execution-time ratio rises from 0.88x at 150k edges, where Souffle is faster, to 5.54x at 1.2M; fused peak device allocations are 85-1,033 MB versus 3,287-44,979 MB for the materializing arm. Exact inference is correctness-equivalent to but slower than ProbLog2. On a public video benchmark, a proximity predicate trained only through symbolic credit replaces hand-set geometry at unchanged held-out accuracy; within Event-Calculus rule search it fails ten-fold cross-validation and does not transfer on a leak-free split. On a maritime corpus, weighted clauses beat crisp selection by 0.065 F1, with the result reproduced by one chronological training pass.

发表机构

  • Brainyblaze Dynamics Inc.(Brainyblaze动力公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑