arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16224cs.CLcs.AI

STAIR:用于时间问答中可解释推理的语义-时间自动机

STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering

Xinlong Dai, Jinchuan Zhang, Lei Gao, Xinzhe Hu, Yuefeng He, Hui Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出STAIR语义-时间自动机,将语义解读与时间推理分离,规则优先设计减少自由推理,在多个时间问答数据集上优于基线,提升了F1值且可解释性更强。

中文摘要 AI 辅助

借助大规模预训练,大语言模型(LLMs)无需针对特定任务训练即可解读各类时间表达式与问题表述。然而,现有的基于提示的神经符号系统仍依赖LLMs完成语义解读与精确时间推理,导致关于时间区间、时间锚点及有序状态的离散决策易受概率误差影响且难以验证。本文提出STAIR,即语义-时间自动机(Semantic-Temporal Automaton for Interpretable Reasoning)。STAIR将语义解读与精确时间推理分离:无答案的LLM适配器将复杂问题表述映射为标准化时间意图,而具有有限控制与防护转换的确定性时间自动机则对规范化证据执行对应策略。遵循规则优先设计,STAIR在规则路径能生成可执行意图时无需调用LLM即可解决标准问题,仅在规则路径失败时应用语义适配。该方法减少了自由形式推理,使时间决策可验证且可解释。具体而言,防护执行支持精确的时间点包含及前后选择,语义适配则处理非精确区间与时间锚定查询。在TimeQA-Easy、TimeQA-Hard、TempReason-L2及TempReason-L3数据集上,STAIR在匹配模型设置的时间问答(TQA)任务中持续优于强基线,使用Qwen2.5-7B模型时平均F1提升16.57%,使用GPT-4o-mini模型时平均F1提升3.10%。此外, ablation( ablation 指消融实验)与诊断分析表明,STAIR在处理边界敏感型与顺序敏感型查询方面表现出色,其防护执行与语义适配分别确保了精确时间点推理与非精确区间处理。

英文摘要

By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro-symbolic systems continue to rely on LLMs for both semantic interpretation and exact temporal inference. Consequently, discrete decisions regarding intervals, time anchors, and ordered states remain vulnerable to probabilistic errors and difficult to verify. We present STAIR, a \textbf{S}emantic-\textbf{T}emporal \textbf{A}utomaton for \textbf{I}nterpretable \textbf{R}easoning. STAIR separates semantic interpretation from precise temporal inference: an answer-free LLM adapter maps complex question formulations to normalized temporal intents, while a deterministic temporal automaton with finite control and guarded transitions executes the corresponding policies over canonicalized evidence. Following a rule-first design, STAIR resolves standard questions without invoking an LLM and applies semantic adaptation only when the rule path fails to produce an executable intent. This approach reduces free-form reasoning, making temporal decisions verifiable and interpretable. Specifically, guarded execution supports precise point-time containment and before/after selection, while semantic adaptation handles non-exact intervals and time-anchored queries. Across the TimeQA-Easy, TimeQA-Hard, TempReason-L2, and TempReason-L3 datasets, STAIR consistently outperforms strong baselines in the TQA task using matched model settings, achieving average F1 improvements of 16.57\% and 3.10\% when utilizing the Qwen2.5-7B and GPT-4o-mini models, respectively. Furthermore, ablations and diagnostic analyses demonstrate that STAIR excels at handling both boundary-sensitive and order-sensitive queries, while its guarded execution and semantic adaptation ensure precise point-time reasoning and inexact intervals, respectively.

↑