arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向表格到多模态报告生成的蒙特卡洛树搜索

Monte Carlo Tree Search for Table-to-Multimodal Report Generation

Teng Lin, Zhiyang Zhang, Yuyu Luo, Nan Tang

arXiv 2608.04071首次发表:更新:

AI 中文总结

针对表格到多模态报告生成的现有方法缺陷,本文提出基于MCTS的MCTS-Report框架,通过分解原子动作与多维奖励函数优化,在自建的MMRBench基准上取得优于基线的77.9总体得分。

AI 中文摘要

从结构化表格数据自动生成包含文本分析和可视化图表的专业多模态报告,是数据智能领域的关键挑战。现有方法采用固定线性流水线与孤立子任务处理,阻碍了事实准确性、视觉质量与叙事连贯性的联合优化。为解决这些问题,本文提出MCTS-Report,一种基于蒙特卡洛树搜索(Monte Carlo Tree Search,MCTS)的框架,将多模态表格到报告生成为结构化搜索空间上的渐进构建过程。核心思路是将报告生成分解为原子动作,包括章节规划、可视化任务识别、图表生成、洞察组织与叙事优化,每个动作由大型语言模型(LLM)基于当前报告状态的动态推理执行。我们在MCTS过程中使用LLM生成逐步推理与动作,在每个节点存储推理轨迹以实现上下文感知的连贯报告构建。为引导搜索,我们设计了多维奖励函数,联合评估数值事实一致性(通过SQL实现)、图表质量、图表-文本对齐性与结构完整性,同时引入多样性惩罚以抑制重复图表,加入前置条件检查以剪枝无效动作。我们还构建了MMRBench,一个包含六个领域真实表格的综合基准,搭配专家优化的参考报告结构与可验证的关键洞察。在MMRBench上的实验表明,MCTS-Report在结构完整性、数值准确性、图表-文本对齐性与洞察新颖性方面显著优于强基线,取得了77.9的总体得分。

英文摘要

Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical challenge in data intelligence. Existing methods suffer from fixed linear pipelines and isolated subtask processing, which hinder joint optimization of factual accuracy, visual quality, and narrative coherence. To address these issues, this paper proposes MCTS-Report, a Monte Carlo Tree Search (MCTS)-driven framework that formulates multimodal table-to-report generation as a progressive construction process over a structured search space. The core idea is to decompose report generation into atomic actions, including chapter planning, visualization task identification, chart generation, insight organization, and narrative refinement, each executed by an LLM based on dynamic reasoning conditioned on the current report state. We use an LLM to generate step-by-step reasoning and actions during MCTS, storing the reasoning trajectory in each node for context-aware, coherent report construction. To guide the search, we design a multi-dimensional reward function that jointly evaluates numerical fact consistency (via SQL), chart quality, chart-text alignment, and structural completeness, while incorporating a diversity penalty to suppress repeated charts and a precondition check to prune invalid actions. We also construct MMRBench, a comprehensive benchmark comprising real-world tables from six domains, paired with expert-refined reference report structures and verifiable key insights. Experiments on MMRBench demonstrate that MCTS-Report significantly outperforms strong baselines across structural completeness, numerical accuracy, chart-text alignment, and insight novelty, achieving a 77.9 overall score.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑