arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33386cs.GRcs.MA

SIVIA-RSI:图解技能的源基适应

SIVIA-RSI: Source-Grounded Adaptation of Diagramming Skills

  • University of Science and Technology of China(中国科学技术大学)
  • Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences(中国科学院苏州生物医学工程技术研究所)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Feng Yuan, Yifan Gao, Haoyue Li, Xin Gao

AI总结:

本文提出SIVIA-RSI框架,通过源基适应实现可复用图解技能迁移,评估显示最强候选在必要关系准确率上优于原始技能,但跨论文迁移一致性不足,强调需源基评估与完整候选覆盖。

AI中文摘要:

科学方法图解通过实体、依赖关系和条件路径来表达计算过程。虽然生成的图形可以通过反复编辑来改进,但尚不清楚一篇论文的经验是否能改善另一篇论文的首张图形。我们提出了SIVIA-RSI,一个用于可复用图解技能的源基适应框架,并通过完整候选评估来研究迁移效果。该框架将批评与源文本段落相关联,对持久技能库提出有界编辑,并将候选竞争与技能接受分离。我们在两篇NLP和语言智能体论文上评估了原始技能和四个学习到的候选技能,每个条件进行两次全新生成。最强候选达到了87.50%的必要关系准确率,而原始技能为83.33%,但候选行为在不同论文间存在差异。在单独开发论文上的局部改进和自动选择器的偏好并未建立一致的迁移效果。将所有22个非正确关系判断追溯到其生成提示,揭示了尽管有明确指令,仍存在不完整的条件规格和歧义。所有十个规划图都使已终止的选定叶子的路径不明确;它们的提示均未明确绑定该路径。我们的发现表明,评估可复用图解技能需要源基关系评估、完整候选覆盖以及检查提示和图像。我们提供了所有二十个迁移输出、技能快照、评估记录和可执行分析。

英文摘要:

Scientific method diagrams express computations through entities, dependencies, and conditional routes. Although generated figures can be improved through repeated editing, it is less clear whether experience from one paper improves the first figure of another. We present \sys, a framework for source-grounded adaptation of reusable diagramming skills, and study transfer through a complete-candidate evaluation. The framework links critiques to source passages, proposes bounded edits to a persistent skill library, and separates candidate competition from skill acceptance. We evaluate the original skill and four learned candidates on two NLP and language-agent papers, with two fresh generations per condition. The strongest candidate attains 87.50\% required-relation accuracy compared with 83.33\% for the original, while candidate behavior differs across papers. Local improvements on a separate development paper and automatic selector preferences do not establish consistent transfer. Tracing all 22 non-correct relation judgments to their production prompts reveals both incomplete conditional specifications and ambiguities despite explicit instructions. All ten planning diagrams leave an already-terminal selected leaf's route unclear; none of their prompts explicitly binds that route. Our findings show why evaluating reusable diagram skills requires source-grounded relation assessment, complete candidate coverage, and inspection of both prompts and images. We provide all twenty transfer outputs, skill snapshots, assessment records, and executable analyses.

↑