发表机构
National Yang Ming Chiao Tung University; University of Illinois Urbana-Champaign(国立阳明交通大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流程图在不同宽高比画布复用时的结构失真问题,提出解析-风格-布局三阶段智能体流程,结合确定性检查与视觉反馈,实现高保真可编辑重排,内容保真度达68.6%。
AI 中文摘要
机器学习论文中的流程图需要在多种画布上复用,包括论文单栏、16:9幻灯片、竖版海报、1:1社交媒体预告图和9:16手机预览。每种格式对同一计算图施加了不同的宽高比,任何静默断开的连接都会错误地表示方法。我们将宽高比自适应流程图重排定义为一个独特任务:给定光栅流程图和目标比例,生成结构保真、无幻觉、可编辑的布局。现有方法表现出特征性失败:图像到图像模型拉伸模块并拒绝极端比例,文本到图像智能体系统产生幻觉内容,解析后渲染系统错误布线。我们提出一个分解为解析、风格和布局阶段的智能体流程,每个阶段将主智能体与批评者配对,批评者结合确定性约束检查和VLM视觉反馈,从而显式检查连通性并防止其被静默破坏。输出为可编辑的mxGraph XML。在包含100个流程图、五种宽高比的精选基准上,由Gemini 3.1 Pro评估并针对人类判断验证,我们的方法达到68.6%的内容保真度,而先前工作为11.2%-41.4%。项目页面:此HTTPS URL
英文摘要
Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are draw.io-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: https://onefigureeverycanvas.vercel.app/
CommentsProject page: https://onefigureeverycanvas.vercel.app/