演化LLM生成特征用于可解释分类
Evolving LLM-Generated Features for Interpretable Classification
浏览论文内容
中文总结 AI 辅助
提出演化框架迭代生成自然语言特征定义,用于可解释分类,在三个基准上平均提升2.9个百分点,优于零样本LLM分类,并提供可审计决策逻辑。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用作分类器,然而它们作为不透明系统运作,其决策难以解释,这使它们在信用评分或医疗诊断等受监管领域的应用复杂化。我们提出一个演化框架,迭代地发现自然语言特征定义(评分标准)以进行可解释分类。LLM生成候选二元特征,对每个样本进行评估,所得向量可用于训练透明分类器(如逻辑回归)。特征集在多次迭代中演化,由分类错误、每类激活率和特征消融分数引导。我们在三个基准上评估,包括一个代表受监管领域的信用风险数据集,比较单次LLM评分标准、演化评分标准和直接零样本LLM分类。演化特征在平均上比单次评分标准提高+2.9个百分点,并在三个任务中的两个上优于零样本LLM分类,同时提供完全可审计的决策逻辑。在信用风险任务上,零样本LLM表现如随机(50.7%),且强烈偏向单一类别,而演化特征实现平衡、可解释的预测。关键的是,这种失败在总体准确性中不可见,仅在逐类审计下显现。分析揭示,当标签边界无法仅从类别名称推断,或LLM缺乏可靠的领域特定推理时,演化最为有效。
英文摘要
Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolutionary framework that iteratively discovers natural language feature definitions (rubrics) for interpretable classification. An LLM generates candidate binary features, evaluates each sample against them, and the resulting vectors can be used to train a transparent classifier such as logistic regression. The feature set evolves over multiple iterations guided by classification errors, per-class activation rates, and feature ablation scores. We evaluate across three benchmarks, including a credit risk dataset representative of regulated domains, comparing single-shot LLM rubrics, evolved rubrics, and direct zero-shot LLM classification. Evolved features improve over single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, while providing fully auditable decision logic. On the credit risk task, the zero-shot LLM performs at chance (50.7%) with a strong bias toward a single class, whereas evolved features achieve balanced, interpretable predictions. Crucially, this failure is invisible in aggregate accuracy and surfaces only under per-class auditing. Analysis reveals that evolution is most effective when label boundaries cannot be inferred from category names alone or when the LLM lacks reliable domain-specific reasoning.
发表机构
- Amazon Web Services(亚马逊云服务)
机构由 AI 辅助整理,请以论文原文为准。