arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31607stat.MEcs.AIcs.LGmath.STstat.TH

黑盒生成式AI的统计属性对齐:通过输出后处理实现

Statistical attribute alignment for black-box generative AI via output post-processing

Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski

首次发表
浏览论文内容

中文总结 AI 辅助

针对黑盒生成式AI,提出通过输出后处理算法最小化查询次数,实现生成输出属性分布与用户目标的对齐,并在文生图等任务上验证了有效性。

中文摘要 AI 辅助

生成式AI系统日益普及,但将其输出与用户需求对齐仍是一个持续的挑战。在此,我们旨在确保AI生成输出的某个属性的分布与用户指定的目标分布对齐。这一需求源于公平性等实例,例如我们希望确保受保护属性(如性别、种族或年龄类别)遵循期望的分布;以及合成数据生成,我们希望生成的数据能代表目标分布。我们研究了实际中重要的黑盒访问设置,即用户可反复查询生成式AI模型。目标是返回m≥1个输出,使其联合属性分布尽可能接近该目标分布。针对精确对齐和近似对齐两种情况,我们开发了能最小化对生成器期望查询次数的算法,并进一步证明了当请求输出数量m趋于无穷大时这些算法的最优性。在文生图生成和地理编码人物生成任务上的实验表明,我们的后处理算法改进了统计属性对齐,补充了基于提示的干预措施。

英文摘要

Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.

↑