发表机构
ML Research Labs(ML研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体文档创作系统,发现结构条件作用存在不对称性,结构利于阅读但不利于写作,还揭示了信息可用性对系统性能的关键影响。
AI 中文摘要
多智能体流水线创作正式文档时,既需读取请求者的表单,也需按要求进行写作。本文报告了一个已部署的投标响应系统,该系统在主权约束下运行开放权重模型,并将其与同一组织实际提交的人工投标进行评估。在系统无可用示例的盲评中,LLM 评判员对其回答的评分:在 55 个真实部分中,至少与人工提交的回答一样好的有 40 个,更好的有 4 个,未出现缺失部分,且总共标记了 1 个无依据的主张。对评判员识别出的每个差距进行分类显示:68%的差距源于系统自身来源中不存在的内容——这是人工作者掌握但从未提供给流水线的知识,因此 15 个不利裁决中仅有 6 个涉及系统本可避免的缺陷。与真实值的偏差更多是信息可用性的结果,而非写作质量的结果,未区分两者的评估会低估此类系统。在此背景下,本文报告了一种条件作用不对称性:将文档呈现为结构标记而非纯散文可提升抽取效果,这在三项阅读任务中得到验证;但该益处无法迁移至条件作用:在配对比较中,将投标的指令材料从散文转换为嵌套 XML,使回答质量从 74%降至 48%。进一步发现:命名禁止性构造会集中而非消除缺陷——96%的剩余缺陷属于提示明确命名的两种形式;将随机注释与确定性窗口函数耦合,可使字节相同文件的抽取需求数量从 68 变为 51。结构应置于模型读取的环节,散文与自应用测试应置于模型写作的环节。
英文摘要
Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-response system, running an open-weights model under sovereignty constraints, and evaluate it against human-written bids the same organisation actually submitted. On a blind comparison where the system had no worked example available, an LLM judge rated its answers at least as good as the human-submitted answer on $40$ of $55$ ground-truth sections, better on $4$, missing on none, and flagged one unsupported claim in total. Classifying every gap the judge identified shows that $68\%$ were content absent from the system's own sources -- knowledge the human author held and the pipeline was never given -- so only $6$ of the $15$ adverse verdicts involve a deficiency the system could have avoided. A divergence from ground truth is more often an information-availability result than a writing-quality one, and evaluations that do not separate the two understate such systems. Against this backdrop we report a conditioning asymmetry. It is well established that rendering documents as structural markup rather than flat prose improves extraction, and we reproduce that on three reading tasks. The benefit does not transfer to conditioning: converting a bid's \emph{instruction} material from prose to nested XML dropped answer quality from $74\%$ to $48\%$ under a paired comparison. We further find that naming a forbidden construction concentrates rather than removes it -- $96\%$ of surviving defects fall in the two forms the prompt explicitly names -- and that coupling a stochastic annotation to a deterministic windowing function moves the extracted requirement count from $68$ to $51$ on a byte-identical file. Structure belongs where the model reads; prose and self-applied tests belong where it writes.
Comments10 pages, 3 figures