arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01813cs.HC

超越指令驱动编辑:面向科学海报的基于来源的问题发现与用户主导修复

Beyond Instruction-Driven Editing: Source-Grounded Problem Discovery with User-Governed Repair for Scientific Posters

  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingda Lyu, Honglin Lu, Xinye Luo, Shiqi Yang

AI总结:

该研究针对交互式编辑器的“表达缺口”问题,提出PROS方法与PROS-Bench数据集,实现科学海报的基于来源的问题发现与用户主导修复,相关评估验证了其有效性并强调需分别评估问题发现等环节。

AI中文摘要:

交互式编辑器通常假设用户已经知道要修改什么内容,但存在一个更早的重要交互状态:用户可能意识到某个作品存在问题,却不知道需要请求什么干预,我们将此称为“表达缺口”。我们推出PROS(Proactive Refinement Of Scientific Posters,科学海报的主动优化),该方法将认知主动性与行为权限分离:系统可呈现基于来源的候选问题,而用户决定哪些问题成为修复目标,以及是否提交最终修改。已接受的问题会移交至原生对象PPTX编辑,并附带验证与可逆预览。我们还推出PROS-Bench,这是一个关联来源的数据集,包含120篇论文和320张可编辑PPTX海报,其中包括120张匹配的核心海报,以及一个独立的会议呈现挑战子集。在核心海报上,PROS在0-100分的VL M(视觉语言模型)评分中,实现了阶段平衡诊断质量的平均得分为67.2;在已接受的诊断中,操作员验证的目标分辨率为87.6%。采用时间盲法自动评分时,已接受目标的论文宏观指标提升了22.7个百分点,但14.8%的可评估已接受目标被拒绝。这种差异表明,问题发现、局部解决方案和实际结果应分别进行评估。更广泛而言,智能编辑器可在用户提出具体编辑请求前支持问题发现,且无需掌控关键变更的权限。

英文摘要:

Interactive editors usually assume that users already know what to change. Yet an important interaction state comes earlier: a user may recognize that an artifact is not working without knowing what intervention to request. We call this the articulation gap. We introduce PROS (Proactive Refinement Of Scientific Posters), which separates epistemic initiative from behavioral authority: the system can surface source-grounded candidate problems, while users decide which become repair goals and whether resulting changes are committed. Accepted issues hand off to native-object PPTX editing with validation and reversible preview. We also introduce PROS-Bench, a source-linked collection of 120 papers and 320 editable PPTX posters, including a 120-poster matched primary core and a separate conference representation challenge. On the primary core, PROS achieves a mean VLM-rated stage-balanced diagnosis quality score of 67.2 on a 0-100 scale and 87.6% operator-verified target resolution among accepted diagnoses. Temporally blinded automated scoring yields a +22.7-point paper-macro accepted-target uplift, yet 14.8% of assessable accepted targets decline. This divergence shows why problem discovery, local resolution, and realized outcome should be evaluated separately. More broadly, intelligent editors can support problem discovery before a concrete edit request exists without taking authority over consequential change.

↑