一种用于LLM文本生成中可靠约束满足的闭环控制架构
A Closed-Loop Control Architecture for Reliable Constraint Satisfaction in LLM Text Generation
- Tampere University(坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种五阶段闭环控制架构,通过确定性代码检查编辑内容,使LLM文本生成在字数或可读性等数值约束上达到92.5%-98.8%的满足率,显著优于单次提示。
AI中文摘要:
软件系统越来越多地将大型语言模型嵌入到必须满足数值输出约束的功能中,即可以用数字或区间表示且可由代码检查的需求,例如目标字数或目标可读性等级区间。由于此类模型是非确定性的,通过自然语言指令而非类型化接口配置,且仅近似满足所述需求,因此单一提示既不能可靠地达到目标,也不能保留源内容。本文提出并评估了一种针对该问题的闭环控制架构。该架构包含五个阶段:生成、评估、调整、存档和分析。模型仅被用于编写和编辑文本,而确定性代码将复合可读性值与目标区间进行比较,拒绝任何删除源实体、数字或关键词的编辑,并做出每个接受决策。在四个商业模型上进行的114次单次生成任务和240次闭环运行中,单次提示在21.1%至31.6%的情况下达到目标,而闭环在92.5%至98.8%的情况下达到目标,平均在两个编辑轮次内,且基于召回率的保真度为0.92至0.93;两种设置共有的两个模型显示出相同的效果。由于控制器优化了用于评判成功度的指标,该结果确立了对已声明、可计算指标的可复现控制,而非经过验证的人类难度。可迁移的实践是将接受条件声明为代码,将模型限制为局部编辑,并对每次编辑进行内容检查。
英文摘要:
Software systems increasingly embed a large language model in features that must satisfy a numeric output constraint, that is, a requirement expressible as a number or an interval and checkable by code, such as a target word count or a target readability grade band. Because such a model is non-deterministic, is configured through natural-language instructions rather than a typed interface, and satisfies a stated requirement only approximately, a single prompt neither reliably meets the target nor preserves the source content. This paper presents and evaluates a closed-loop control architecture for this problem. It has five stages: generate, evaluate, adjust, archive, and analyze. The model is called only to write and to edit text, while deterministic code compares a composite readability value against a target band, rejects any edit that drops source entities, numbers, or keywords, and makes every accept decision. Over 114 single-shot generation jobs and 240 closed-loop runs on four commercial models, single-shot prompting met the target in 21.1 to 31.6 percent of cases and the closed loop in 92.5 to 98.8 percent, within two edit rounds on average and at a recall-based fidelity of 0.92 to 0.93; the two models common to both settings show the same effect. Because the controller optimizes the value on which success is scored, the result establishes reproducible control over a declared, computable metric and not validated human difficulty. The transferable practice is to declare the acceptance condition as code, bound the model to local edits, and gate every edit on a content check.