arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13624cs.CLcs.AIcs.SD

基于语义感知偏差估计的大音频语言模型公平性测量

Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

Zhe Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大音频语言模型公平性评估中语义与说话人混淆因素的问题,提出语义感知混合效应回归框架,经实验验证可减少虚假结论,得到更稳健的子群体性能差异估计。

中文摘要 AI 辅助

大音频语言模型(Large Audio Language Models, LALMs)在语音识别、音频问答等音频理解任务中的应用日益广泛,引发了对不同人口统计子群体公平性的关注。口语输入场景下的公平性评估颇具挑战,存在诸多混淆因素,包括口语内容的语义差异和说话人特定特征,忽略这些因素可能会得出关于模型偏差的误导性结论。本文提出一种用于LALMs公平性评估的语义感知混合效应回归框架,明确考虑上述混淆因素。该方法将参考文本的句子级语义嵌入作为协变量,并将说话人身份建模为随机效应;值得注意的是,语义表征是从待评估的同一LALM中提取的,可实现对模型自身感知到的差异进行语义控制。在模拟数据和真实基准上开展的实验表明,所提方法大幅减少了虚假公平性结论,能得出更稳健、可解释的子群体性能差异估计值。

英文摘要

Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation in spoken-input settings is challenging due to confounding factors, including semantic variation in spoken content and speaker-specific characteristics. Ignoring these factors can result in misleading conclusions about model bias. We propose a semantic-aware mixed-effects regression framework for fairness evaluation in LALMs that explicitly accounts for these confounders. Our approach incorporates sentence-level semantic embeddings of reference text as covariates and models speaker identity as a random effect. Notably, semantic representations are extracted from the same LALM under evaluation, enabling semantic control over variation as perceived by the model itself. Experiments on simulated data and real-world benchmarks demonstrate that the proposed approach substantially reduces spurious fairness findings and yields more robust and interpretable estimates of subgroup performance differences.

发表机构

  • Meta Platforms, Inc.(元平台公司)

机构由 AI 辅助整理,请以论文原文为准。

↑