发表机构
University of Cambridge; Chinese Academy of Sciences; University of Chinese Academy of Sciences(剑桥大学; 中国科学院; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对从发明人风格披露文件生成专利申请的现实需求,构建了Dis2Pat数据集,并提出可本地部署的多智能体框架Patent-MAF,实验表明该框架在专利撰写任务上表现优于多数开源模型且可与闭源模型竞争。
AI 中文摘要
尽管近期大型语言模型(LLMs)在单个专利撰写任务上已取得令人满意的结果,但它们根本无法解决现实世界专利撰写的核心挑战:从早期发明材料直接生成完整且符合法律要求的专利申请。现有研究主要假设输入是后期阶段、高度结构化或已符合法律规范的内容,然而实际专利流程始于发明人撰写的非正式、非法律化的披露文件。为弥合这一差距,我们推出Dis2Pat——一个披露转专利数据集,其通过要求直接从发明人风格的非法律化披露文件生成完整专利申请,来反映现实的专利流程。考虑到长篇幅、受法律约束的专利撰写固有的难度以及严格的隐私要求,我们进一步提出名为Patent-MAF的强基线模型,它是一个可本地部署的多智能体专利撰写框架。基准测试结果显示,当前LLMs在专利撰写方面存在局限,而Patent-MAF作为强基线模型,始终优于被评估的开源模型,且与大型闭源模型相比具有竞争力。
英文摘要
While recent large language models (LLMs) have achieved promising results on individual patent drafting tasks, they fundamentally fail to investigate the core challenge of real-world patent drafting: generating a complete and legally coherent patent application directly from early-stage invention materials. Prior work predominantly assumes later-stage, highly structured, or already legalistic inputs. However, real patenting workflows begin with informal, de-legalized disclosures authored by inventors. To bridge the gap, we introduce Dis2Pat, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures. Given the inherent difficulty of long-form, legally constrained patent drafting and the strong privacy requirements, we further propose a strong baseline named Patent-MAF. It is a multi-agent framework for locally deployable patent drafting. Benchmark results reveal that current LLMs exhibit limitations in patent drafting, while Patent-MAF provides a strong baseline that consistently outperforms evaluated open-source models and remains competitive with large closed-source models.
CommentsAccepted to EMNLP 2026