arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PatchWrite:一行,而非一节——面向AI起草手稿的编译门控、有效性保留编辑

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

Weiwei Yang

arXiv 2608.23001首次发表:更新:

发表机构

Solus(索勒斯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI手稿编辑易改动无关内容的问题,PatchWrite通过编译门控与证据锁机制,在压力测试与盲评中均展现出更好的事实保留效果,且能有效修复注入的缺陷。

AI 中文摘要

自动化手稿流水线常通过重新生成整个章节来修复局部缺陷,即便生成的PDF仍可编译,也会导致不相关的指标和参考文献发生改变。PatchWrite则对候选编辑如何提交为手稿状态进行了约束:它复用了有限的EDIT N M编辑与回退机制,但通过致命日志检查收紧了编译接受条件,并新增了证据锁,要求每个引用的关键项和实验数值标记都需经过参考注册表或实验日志的验证。未通过任一检查的候选编辑将被拒绝,保留之前的HEAD状态。在24份手稿×8个缺陷的预言机压力测试中,共768个任务,编译中断缺陷与仅内容缺陷各占一半,全槽重写会在所有案例中修改不相关的“12层”行,192个案例中0个被保留,数值Jaccard指数为0.6667,而PatchWrite在192个案例中全部保留了该行。移除编译门将接受率降至0,移除证据锁则会让虚构的参考文献通过,所有8个缺陷均呈现此模式。为用生成而非预言机编辑测试该协议,我们让生成模型提出编辑后重新运行192个任务,模型的候选编辑在75%的案例中被接受,几乎所有拒绝都来自一种可复现的失败模式:模型试图删除一行时使用了当前语法不支持的空替换。所有被接受的候选编辑均通过了两道检查,其中93.75%修复了注入的缺陷,其余案例涉及技术上有效但句子不合适的参考文献,以及一次标记变更的差一点成功。在16份PDF对的盲评中,两名评估者均因PatchWrite保留了实验室验证的事实而更偏好它,C1李克特评分分别为5.0和2.0,而对散文质量的评分几乎相同。193项产品内起草任务的日志显示,实际中也存在同类失败。

英文摘要

Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing unrelated metrics and citations to change even when the resulting PDF still builds. PatchWrite instead constrains how candidate edits become committed manuscript states: it reuses bounded EDIT N M editing and rollback, but tightens compilation acceptance with fatal-log checks and adds evidence locks that require every cited key and experimental numeric token to be attested by a reference registry or experimental log. Candidates that fail either check are rejected and the previous HEAD is retained. On a 24-manuscript x 8-fault oracle stress test (768 jobs, evenly split between compile-breaking and content-only faults), whole-slot rewriting mutated an unrelated "12-layer" line in every case (0/192 preserved; numeric Jaccard 0.6667), whereas PatchWrite preserved it in 192/192 cases. Removing the compile gate reduced acceptance to 0, while removing the evidence gate allowed a hallucinated citation to pass. The same pattern held across all eight faults. To test the protocol with generation rather than oracle edits, we reran the 192 jobs with the writer model proposing the edits. The model's candidates were accepted in 75% of cases; nearly all rejections came from one reproducible failure mode in which the model attempted to delete a line using an empty replacement unsupported by the current grammar. Every accepted candidate passed both gates, and 93.75% fixed the injected fault; the remaining cases involved a technically valid but sentence-inappropriate citation and one markup-changing near-miss. In a blind evaluation of sixteen PDF pairs, both raters preferred PatchWrite for preserving lab-grounded facts (C1 Likert 5.0 vs. 2.0), while rating prose quality nearly identically. Logs from 193 in-product drafting tasks show the same classes of failures occurring in practice.

Comments12 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑