发表机构
Jiutian(九天)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TrustDABench基准,评估8种LLM在结构化数据分析中的可靠性与鲁棒性,发现现有模型在检测冲突证据、应对扰动等方面存在不足,需提升证据边界识别与不变性推理能力。
AI 中文摘要
大语言模型(LLMs)正越来越多地被用于分析电子表格、CSV文件及其他结构化数据,但生成看似正确的答案与生成可信的分析并不等同。可信的结果应当具备从用户问题到相关数据证据的有效路径支撑。这一要求引出两个诊断性问题:当该路径不存在时,LLM能否拒绝回答或要求澄清;以及当相同证据以不同表格形式呈现时,它能否保持正确的分析结果。我们提出TrustDABench这一基准,将上述问题具体化为可靠性与鲁棒性。从证据-路径视角出发,我们推导了19种扰动算子,并通过基于Agentic-LLM的生成框架实现这些算子。TrustDABench包含2340个人工验证的扰动实例,我们对8种代表性LLM进行了评估。结果显示存在显著提升空间:最佳可靠性结果是GPT-5.5达到的平均MRS仅为24.21%,最佳鲁棒性结果是Claude-Sonnet-5达到的平均ASR仍有9.10%。失败是系统性的:模型很少检测到冲突证据,常沿着可执行但无支撑的分析路径继续,且对改变观测边界或跨表关系的扰动保持敏感。这些发现表明,可靠的结构化数据分析仍需更强的证据边界识别与表示不变性推理能力。
英文摘要
LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis. A trustworthy result should be supported by a valid path from the user question to the relevant data evidence. This requirement creates two diagnostic questions: whether an LLM can refuse to answer or ask for clarification when such a path does not exist, and whether it can preserve the correct analysis when the same evidence is expressed in different table forms. We introduce TrustDABench, a benchmark that operationalizes these questions as reliability and robustness. Starting from the evidence-path view, we derive 19 perturbation operators and instantiate them through an Agentic-LLM-based generation framework. TrustDABench contains 2,340 human-verified perturbed instances, and we evaluate eight representative LLMs. The results show substantial headroom: the best reliability result is only 24.21% average MRS, achieved by GPT-5.5, while the best robustness result still has 9.10% average ASR, achieved by Claude-Sonnet-5. The failures are systematic: models rarely detect conflicting evidence, often continue along executable but unsupported analysis paths, and remain sensitive to perturbations that change observation boundaries or cross-table relations. These findings suggest that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.
CommentsCode&Data: https://github.com/Skyorca/TrustDABench