arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05686cs.AI

时间序列问答系统是否读取了时间序列?证据使用与推理可靠性

Do Time-Series QA Systems Read the Time Series? Evidence Use and Reasoning Reliability

Zhuomin Chen, Jingchao Ni, Xu Zheng, Janki Bhimani, Mo Sha, Wei Cheng, Dongsheng Luo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估四个时间序列问答系统,引入COMMON-TSQA基准及干预测试,发现聚合性能掩盖证据使用问题,且推理依据常缺乏事实基础。

中文摘要 AI 辅助

近年来,时间序列问答(QA)系统取得了显著进展。然而,生成正确答案并不能表明保留所提供的数值序列是否改善了任务性能,也不能表明预测是否对该输入的变化敏感。虽然一些系统提供了推理依据,但答案准确性也不能表明其数值主张是否基于所提供的序列,或所陈述的推理是否有效。在这项工作中,我们重点评估四个时间序列问答系统:TimeOmni-1、ChatTS、TimeOmni-VL和Time-MQA。首先,对于三个发布了评估数据的系统,我们复现其报告的结果,并将系统与其骨干模型的性能进行比较。然后,我们引入一个名为COMMON-TSQA的基准,它从现有时间序列基准中收集公开评估数据集,并统一其样本表示、任务定义和答案模式,同时在共同评估标准下通过每个系统自身的接口进行评估。评估使用原始条件和六种干预,同时保持问题和目标不变。我们的分析表明,仅聚合性能可能掩盖系统如何使用数值证据。尽管个别预测发生显著变化,但相似的任务级分数仍可能出现。某些干预会引发简单的回退行为,而非保持任务能力。我们还评估了推理依据的事实基础、推理有效性以及与最终答案的一致性。我们发现,推理依据常常包含输入不支持的时间序列主张。此外,推理依据审计表明,推理依据与其最终答案之间的一致可能与错误的数值描述或无效的中间推理共存。

英文摘要

In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.

发表机构

  • Florida International University(佛罗里达国际大学)
  • University of Houston(休斯顿大学)
  • NEC Laboratories America(NEC美国实验室)
  • Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

↑