AI 中文总结
研究生成模型引导问题,发现训练数据设定了属性引导预算,分为旋钮可触及和示例可触及两部分。提出用示例引导的方法,经审核训练数据构建示例集,能带来范围覆盖全预算及可引导至特定目标的好处,并在相关领域验证。
AI 中文摘要
生成模型通过旋钮(如提示、引导尺度、属性标签)进行引导。但过了某个点,旋钮转动对所关注属性不再起作用。我们发现这并非模型缺陷,而是由训练数据设定的预算。属性可移动范围分为两部分,一部分旋钮能触及,另一部分只有示例能触及。通常后者范围大得多。获取第二部分需展示示例而非转动旋钮。通过廉价审核训练数据可衡量预算,我们给出构建能触及全部预算的示例集的方法。这带来两点好处:范围覆盖整个预算;能引导至只能通过示例指定的目标。我们将其转化为可证伪的论断,并在图像和晶体结构生成两个不相关领域验证,明确旋钮和示例各自适用之处。
英文摘要
Generative models are steered with knobs -- prompts, guidance scales, property tags. Turn one as hard as you like and, past a point, it stops moving the property you care about. We find that ceiling is not a shortcoming of the model but a budget, set by the training data before the model is trained: a property's movable range splits in two -- the part a knob can reach, and a second, significant part that only examples -- concrete instances of what you want more of -- can reach. That second part is usually much larger, but not always, and the same budget says so in advance. Reaching that second part takes a different move: instead of turning a knob, you show the model examples, composed from what it already learned rather than added to its training. A cheap audit of the training data measures the budget; we give a recipe for building the example set that reaches all of it. This buys two things a knob can't. Reach: it moves a property across the whole budget, not just the part a knob reaches. Expressiveness: it steers toward targets you can only specify by example -- including ones you can't put into words. We turn these into a handful of falsifiable claims and verify them in two unrelated domains, image and crystal-structure generation -- marking where a knob is enough, and where only examples will do.
Comments39 pages, 14 figures