arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24562cs.AI

语言模型中用于选择性预测的分层组条件共形风险控制

Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

Murilo Salem, Luísa Böhm, Daniel Pontes, Anderson Ferrugem

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对语言模型服务异质群体时共形风险控制的问题,提出HG-CRC框架,通过特定策略在组层次结构节点上保证风险,经实验评估,该框架在ARC挑战等任务中对高精度模型效果良好,不同基准结果有差异,分层深度对消除预算问题有作用。

中文摘要 AI 辅助

大型语言模型服务于由领域、主题难度和语言风格构成的异质群体。共形风险控制(CRC)为带有弃权的选择性预测提供了严格的边际风险保证,但边际保证并不意味着每个组的保证:一个模型可以满足总体预算,却系统性地使子群体过度暴露于错误中。在组构成存在轻微变化时,标准CRC在高达47%的试验中违反预算。我们提出了HG-CRC(分层组条件CRC),这是一个事后校准框架,可在用户定义的组层次结构的所有节点上强制执行同时的风险保证。它对节点应用邦费罗尼校正,并采用叶优先策略,即使用最具体适用的阈值,当更精细的节点未通过认证或拒绝示例时,退回到更粗粒度的节点。它只需要一个留出的校准集,无需重新训练。我们在三个模型(Qwen3-4B、Llama-3.1-8B-Instruct、Gemma-3-4B)和两个基准(ARC挑战、MMLU-Pro)上进行了评估,涵盖了八种配置,以探究独立同分布泛化、异质性、混合/领域/提示/难度转移、标签噪声和量化。主要结果:对于高精度模型(Qwen3-4B、Llama-3.1-8B),HG-CRC在ARC挑战上的经验违反率为0%且加权预期风险(WGER)=0。在500次自助试验中,这些零是经验上限(真实率高达0.6%),未得到认证。结果因基准而异:在MMLU-Pro上,这些模型要么完全弃权,要么(Llama)保持WGER=0.014。Gemma-3-4B在此处校准不佳,通过弃权而优雅地降级。参与成本与全局CRC相比为22至37分。消融实验表明分层深度可消除预算问题:去除难度级别后,违反率回升至约11%。理论保证需要邦费罗尼校正,不过其经验效果仅在有许多节点时才重要。

英文摘要

Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees do not imply per-group ones: a model can meet the population budget while systematically over-exposing subgroups to errors. Under mild shift in group composition, standard CRC violates the budget in up to 47% of trials. We propose HG-CRC (Hierarchical Group-Conditional CRC), a post-hoc calibration framework enforcing simultaneous risk guarantees across all nodes of a user-defined group hierarchy. It applies a Bonferroni correction over nodes and a leaf-first policy that uses the most specific applicable threshold, falling back to coarser nodes when a finer one is uncertified or rejects the example. It needs only a held-out calibration set, with no retraining. We evaluate on three models (Qwen3-4B, Llama-3.1-8B-Instruct, Gemma-3-4B) and two benchmarks (ARC Challenge, MMLU-Pro) across eight configurations probing IID generalization, heterogeneity, mixture/domain/prompt/difficulty shift, label noise, and quantization. Main result: HG-CRC reaches an empirical 0% violation rate and WGER=0 on ARC Challenge for high-accuracy models (Qwen3-4B, Llama-3.1-8B). At 500 bootstrap trials these zeros are empirical upper bounds (true rate up to 0.6%), not certified. Results are benchmark-specific: on MMLU-Pro these models abstain entirely or (Llama) retain WGER=0.014. Gemma-3-4B, poorly calibrated here, degrades gracefully by abstaining. Participation cost vs. global CRC is 22 to 37 points. Ablations show hierarchical depth clears the budget: removing difficulty level returns violations to about 11%. Bonferroni is needed for the theoretical guarantee, though its empirical effect matters only with many nodes.

↑