PosterHarness:将科学海报生成转变为可审计的指令跟随基准
PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark
浏览论文内容
中文总结 AI 辅助
研究针对文本丰富图像模型,缺乏衡量其是否遵循科学传播规范的方法这一问题,提出PosterHarness,将海报生成变为可审计任务,介绍核心方法,展示了相关实验发现及与其他方法的对比结果。
中文摘要 AI 辅助
富含文本的图像模型如今能设计海报规模的布局,但缺乏衡量其是否遵守科学传播契约的方法。我们提出POSTERHARNESS,一个可审计的框架,将海报生成重新定义为可衡量的指令跟随任务,带有试点基准和失败分类法。POSTERHARNESS使用占位符优先契约来分离模型通常混淆的两项工作。模型执行视觉摘要设计,排版、阅读路径、颜色和背景,但从不绘制带有数据的图形。每个图形区域必须是一个空的带标签占位符;一个确定性合成器在检测到的坐标处插入真实的源论文图形。这使得属性可衡量,占位符数量和ID准确性、空白度、宽高比合规性、不使用合成图形、公共文本卫生以及源图形出处,失败记录为明确拒绝,而非隐藏在看似合理的输出中。我们在12篇论文上实例化该框架并报告三项发现。(i)一个反事实探针显示占位符契约使三篇论文中基于VLM计数的合成图形从34降至0。(ii)一个失败分类法识别出阻碍契约:占位符几何形状、占位符质量保证、模板批评家和公共文本。(iii)与Paper2Poster的比较显示了一种权衡:PosterHarness产生更高分辨率的工件、更低的白色画布比例和更强的VLM视觉偏好;确定性基线保留了略多的PosterQuiz风格信息且运行更快。我们将此报告为机制特征描述,而非优越性声明。所有工件、提示、清单和审计脚本都作为可重复使用的评估组件发布。
英文摘要
Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures. We present POSTERHARNESS, an auditable harness reframing poster generation as measurable instruction-following tasks, with a pilot benchmark and failure taxonomy. POSTERHARNESS uses a placeholder-first contract to separate two jobs models otherwise conflate. The model performs visual-summary design: typography, reading path, color, and background -- but never draws data-bearing figures. Every figure region must be an empty labeled placeholder; a deterministic compositor inserts real source-paper figures at detected coordinates. This makes properties measurable: placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance -- with failures logged as explicit rejections, not hidden in plausible-looking output. We instantiate the harness on 12 papers (6 HEP, 6 AI/ML-adjacent) and report three findings. (i) A counterfactual probe shows the placeholder contract drives VLM-counted synthesized figures from 34 to 0 across three papers. (ii) A failure taxonomy identifies blocking contracts: placeholder geometry, placeholder QA, template critic, and public text. (iii) Comparison with Paper2Poster shows a trade-off: PosterHarness yields higher-resolution artifacts, lower white-canvas fraction, and stronger VLM visual preference; the deterministic baseline retains slightly more PosterQuiz-style information and runs faster. We report this as regime characterization, not a superiority claim. All artifacts, prompts, manifests, and audit scripts are released as a reusable evaluation component.
发表机构
- School of Physics, Peking University(北京大学物理学院)
机构由 AI 辅助整理,请以论文原文为准。