发表机构
Systems Research Institute, Polish Academy of Sciences(波兰科学院系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于单元格角色标注的电子表格分块框架,提升RAG可解释性,并指出需通过降维技术将二维表格展平为一维文本以突破瓶颈。
AI 中文摘要
语义单元格标注提升了LLM驱动的RAG系统中电子表格分块的可解释性,通过丰富上下文而非提高检索准确性来辅助答案生成。我们提出了一种新框架,利用单元格角色标注将任意电子表格分割为可解释的分块。我们的框架超越了现有技术水平,但面临一个硬性上限。电子表格本质上是二维非结构化数据,具有连续关系和无限潜在的单元格角色。由于分类模型局限于有限的、预定义的类别,它们无法完美捕捉这种结构细微差别,即使达到人类水平的标注也是如此。我们表明,解决电子表格到LLM的瓶颈需要超越离散单元格分类。相反,该领域必须开发降维技术,直接将二维非结构化电子表格展平为一维非结构化文本。文本分块将更易于下游RAG解释和生成。
英文摘要
Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework beats the state of the art, yet it faces a hard ceiling. Spreadsheets are fundamentally two-dimensional unstructured data with continuous relationships and infinite potential cell roles. Because classification models are restricted to finite, pre-defined classes, they cannot perfectly capture this structural nuance, even with human-level annotation. We show that addressing the spreadsheet-to-LLM bottleneck requires moving beyond discrete cell classification. Instead, the field must develop dimensionality-reduction techniques to directly flatten 2D unstructured spreadsheets into 1D unstructured text. Text chunks would be easier for downstream RAG to interpret and generate from.