arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习结构收敛:用于时间推理的神经符号基准测试

TRACTA: Benchmarking Temporal Reasoning over Semantic Trajectories

Michael Romei De Socio, Gian Luca Pozzato, Alessio Merlo

arXiv 2607.22365首次发表:更新:

发表机构

University of Turin; CASD – School of Advanced Defense Studies(都灵大学; 高级国防研究学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍用于高复杂性事件驱动系统时间结构推理的TRACTA基准测试,含三个任务,比较多种模型。结果显示基于语义轨迹的时间建模效果最佳,消融与捷径诊断分析揭示相关信息,表明语义基础轨迹对时间结构推理有效,支持语义接口研究。

AI 中文摘要

高复杂性操作环境需要能检测和预测时间分布模式而非对孤立事件进行分类的方法。本文介绍了TRACTA(时间推理与能力-轨迹分析),这是一个通过类似多域操作(MDO)场景实例化的、用于高复杂性事件驱动系统中时间结构推理的受控合成基准测试。该基准测试包括三个任务,并比较了原始事件神经模型、轻契约语义基线和基于语义基础轨迹运行的神经符号配置。结果表明原始事件级学习仍有信息价值,但基于语义能力和上下文直接影响轨迹的时间建模获得了最高的总点估计,消融分析表明能力动态、上下文影响和时间结构贡献了互补信息,捷径诊断表明在主要神经输入视图中最直接的跨运行全局标识符捷径得到了控制。总体而言,研究结果支持一个有限的方法结论:在受控合成环境中,语义基础轨迹为时间结构推理提供了有效表示,支持对事件数据、结构化表示和时间学习之间语义接口的进一步研究。

英文摘要

High-complexity operational environments require methods that characterize temporally distributed patterns rather than classify isolated events. This paper introduces TRACTA (Temporal Reasoning and Capability-Trajectory Analysis), a knowledge-aligned synthetic benchmark for temporal structural reasoning, instantiated through Multi-Domain Operations (MDO)-like scenarios. TRACTA defines offline structural annotations over contextual direct-impact and accumulated capability trajectories and evaluates three tasks: early_warning, pattern_detection, and run_classification. The frozen comparison includes raw-event neural references, a contract-lite rule comparator, and a recurrent semantic-input reference. The semantic-input recurrent reference has the highest aggregate macro-F1 point estimates, with the largest margins on the two temporal tasks, while raw-event references remain predictive and lead in four individual early-warning target--lead settings. Component-zeroing diagnostics show that both semantic trajectory blocks contain useful signal within the evaluated recurrent configuration. Run-local aliasing removes stable cross-run target and location identities from the primary raw input, although executed diagnostics retain shallow predictivity. These results are configuration-level: semantic inputs are aligned with the benchmark's target-generation space, and the evaluated systems also differ in architecture, training, and available information. TRACTA therefore provides a reproducible testbed for examining knowledge-aligned temporal prediction, not evidence of a causal representation advantage, statistically resolved superiority, or operational readiness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑