arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

任意电子表格问答需要解读其网格结构

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Zofia Smoleń

arXiv 2609.20732首次发表:更新:

发表机构

Systems Research Institute, Polish Academy of Sciences(波兰科学院系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于单元格角色标注的电子表格分块框架,提升RAG可解释性,并指出需通过降维技术将二维表格展平为一维文本以突破瓶颈。

AI 中文摘要

语义单元格标注提升了LLM驱动的RAG系统中电子表格分块的可解释性,通过丰富上下文而非提高检索准确性来辅助答案生成。我们提出了一种新框架,利用单元格角色标注将任意电子表格分割为可解释的分块。我们的框架超越了现有技术水平,但面临一个硬性上限。电子表格本质上是二维非结构化数据,具有连续关系和无限潜在的单元格角色。由于分类模型局限于有限的、预定义的类别,它们无法完美捕捉这种结构细微差别,即使达到人类水平的标注也是如此。我们表明,解决电子表格到LLM的瓶颈需要超越离散单元格分类。相反,该领域必须开发降维技术,直接将二维非结构化电子表格展平为一维非结构化文本。文本分块将更易于下游RAG解释和生成。

英文摘要

Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework beats the state of the art, yet it faces a hard ceiling. Spreadsheets are fundamentally two-dimensional unstructured data with continuous relationships and infinite potential cell roles. Because classification models are restricted to finite, pre-defined classes, they cannot perfectly capture this structural nuance, even with human-level annotation. We show that addressing the spreadsheet-to-LLM bottleneck requires moving beyond discrete cell classification. Instead, the field must develop dimensionality-reduction techniques to directly flatten 2D unstructured spreadsheets into 1D unstructured text. Text chunks would be easier for downstream RAG to interpret and generate from.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑