arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

困惑陷阱:专利法何时让人类写作看起来像人工智能

The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI

Anubhab Banerjee

arXiv 2607.13044首次发表:更新:

发表机构

European Patent Office(欧洲专利局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对专利法下筛选疑似人工智能生成专利文本的难题,在消费级硬件条件下用多种策略进行基准测试,发现现有检测器误报率高,问题是结构性的,提出七特征语言复杂性逻辑回归方法,提升了准确率。

AI 中文摘要

欧洲专利局(EPO)报告称2025年专利申请量创纪录,2026年EPO指南要求申请人对第83条和第42条规定下的大语言模型辅助内容严格负责,这给筛选疑似人工智能生成的专利文本带来压力。但存在两个限制因素,一是实际审查环境通常只有约8GB显存的消费级GPU而非数据中心级评分堆栈,二是欧洲专利公约第84条要求权利要求清晰简洁,这使得人类起草的文本与大语言模型处于相同的低困惑度、低突发性流形上。研究在消费级硬件条件下,用五种提示策略对500项已授权的EPO H04电信专利和500项大语言模型生成的对应文本进行基准测试。在权利要求层面,所有检测器的误报率均超过60%。研究认为该问题是结构性的而非替代模型能力问题。一种七特征语言复杂性逻辑回归在28.1%的误报率下达到74.0%的准确率,相比仅使用困惑度的基线在可比操作点上有13个百分点的提升,且推理时不使用似然性,在相同硬件预算内。

英文摘要

The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted content under Article 83 and Rule 42, creating pressure to triage suspected AI-generated patent text. Two constraints make this hard. First, realistic prosecution settings often have only consumer GPUs with about 8 GB VRAM, not datacenter-class scoring stacks. Second, Article 84 of the European Patent Convention requires claims to be clear and concise, pushing human drafting onto the same low-perplexity, low-burstiness manifold that LLMs occupy. We benchmark three open-source zero-shot detectors on 500 granted EPO H04 telecom patents versus 500 LLM-generated counterparts using five prompting strategies, all under the consumer hardware envelope. At claim level, all detectors exceed 60 percent false-positive rate: Binoculars 78.3 percent, Fast-DetectGPT 61.3 percent, DetectGPT 80.5 percent. The failure persists under Qwen2.5-3B-Instruct regeneration, LoRA-adapted Pythia-2.8B scoring heads, cross-IPC replication on A61K, C07D, and F03D (mean FPR 84.6 percent), and H100 re-evaluation with published Falcon-7B and GPT-J-6B heads, arguing the issue is structural rather than substitute-model capacity. A seven-feature linguistic-complexity logistic regression reaches 74.0 percent accuracy at 28.1 percent FPR, a 13 percentage-point gain over a perplexity-only baseline at a comparable operating point, without using likelihood at inference and within the same hardware budget.

Journal refProceedings of the ICML 2026 Workshop on AI4Law

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑