SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models
SPAN:大型语言模型跨日历时间推理的基准测试与改进
机构 * Li Auto(利亚特)
专题命中 代码评测 :code generation(abstract);分类 cs.AI
AI总结 SPAN通过开发时间代理提升大型语言模型跨日历时间推理能力,实验显示其准确率高达95.31%
Comments Accepted at the AAAI 2026 conference. This version includes the supplementary appendix