arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PaperBanana-Interact:基于多轮人类反馈的科学图表优化

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng

arXiv 2608.30241首次发表:更新:

发表机构

Google; Peking University; University of California, Los Angeles; University of Illinois Urbana-Champaign(谷歌公司; 北京大学; 加利福尼亚大学洛杉矶分校; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对科学图表多轮优化的空白,构建了多轮图表生成基准MTPaperBananaBench,提出多智能体系统PaperBanana-Interact,解决基线系统的质量漂移与遗忘问题,显著提升图表优化效果。

AI 中文摘要

近期已有研究尝试从论文内容自动生成科学图表(Lin等,2026;Zhu等,2026a)。然而,仅通过单轮交互完全满足作者的视觉偏好与沟通需求颇具挑战性:在我们的初步用户研究中(样本量N=14),所有参与者在查看初始草稿后均要求进一步修改,其中86%的参与者认为优化后的图表满意度更高。尽管需求明确,但多轮优化流程仍未得到充分探索。为填补这一空白,我们构建了MTPaperBananaBench,这是一个包含292张图片、标注有3518条用户需求的多轮图表生成基准。为减少昂贵的人工研究并实现可扩展的基准测试,我们构建了一个用户模拟器,该模拟器在每轮交互中识别未被满足的需求,并将其中k条转化为自然语言反馈。通过对需求满足度和图表整体质量的评估,我们发现现有多轮基线系统存在两类关键失败模式:(1)质量漂移,即图表质量随轮次推进逐渐下降;(2)遗忘,即已实现的特征在后续轮次中丢失。为解决这些问题,我们提出了PaperBanana-Interact,这是一个通过内部批评-优化循环来优化图表的多智能体系统。PaperBanana-Interact在各轮次中始终提升而非降低图表质量,其质量得分较基线系统高出11.9-18.6个百分点,遗忘率降低3.7-6.2个百分点。

英文摘要

Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.

Commentshttps://shirley-wu.github.io/PaperBanana-Interact/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑