arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22956cs.CL

控制的错觉:为什么单纯的分类器反演在概念瓶颈文本生成中会悄然失效

The Illusion of Control: Why Bare Classifier Inversion Silently Fails in Concept-Bottleneck Text Generation

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Qi Bing, Xiaowei Shao

AI总结:

该研究针对概念瓶颈文本生成,发现单纯分类器反演因生成流形外代码而失效,其性能逊于简单事后先验,还验证了该问题诊断并实现与基线的公平对比。

AI中文摘要:

概念瓶颈可控生成通过低维概念代码实现多属性控制,部署时需从目标属性配置合成该代码。我们在概念瓶颈文本生成的多轴组合泛化场景下研究此问题,对比三种推理时代码获取方式:针对编码器头的分类器反演、参考文本编码、事后标签条件先验。由于概念代码不具备直接的语言模型流畅度项,正则化反演需将代码约束至编码器的训练分布。因此我们测试了单纯反演及三种正则化变体:标签无关与标签相关的马氏距离惩罚、条件归一化流密度基线。在覆盖1.24亿至80亿参数的三种主干模型上,所有测试的反演变体均逊于针对相同检查点按各组合编码器均值拟合的简单事后先验。单纯分类器反演还会悄然崩溃至随机水平,根源在于直接测得的流形外代码。我们在真实世界基准及外部评估器下验证了该诊断,实现了与已发表基线的公平对比。

英文摘要:

Concept-bottleneck controllable generation routes multi-attribute control through a low-dimensional concept code that, at deployment, must be synthesised from a target attribute configuration. We study this problem in concept-bottleneck text generation under multi-axis compositional generalisation, comparing three ways to obtain the inference-time code: classifier inversion against the encoder heads, reference-text encoding, and a post-hoc label-conditioned prior. Since a concept code admits no direct LM-fluency term, regularising inversion must instead constrain the code toward the encoder's training distribution. We therefore test bare inversion and three regularised variants: label-agnostic and label-conditioned Mahalanobis penalties, and a conditional normalising-flow density baseline. Every inversion variant we test underperforms a simple post-hoc prior fitted to per-combination encoder means on the same checkpoints, across three backbone families spanning $124$M to $8$B parameters. The bare form of classifier inversion also silently collapses to chance, traceable to a directly measured off-manifold code. We validate this diagnosis on real-world benchmarks and under external evaluators, enabling fair comparison with published baselines.

补充信息

↑