arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20427cs.CVcs.AI

语言接地解释何时有帮助?用于农场监测可解释绵羊面部疼痛的图瓶颈

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

Alam Noor, Miguel Guti'errez Gait'an

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过概念瓶颈架构,证明语言接地解释在绵羊面部疼痛识别中不可靠,提出仅用SPFES概念分数的方法,牺牲少量一致性换取可学习概念,并揭示架构必要性不等于语义有效性。

中文摘要 AI 辅助

基于面部表情的自动疼痛识别可以使绵羊的持续福利评估变得实用,但采用取决于信任:饲养员无法对没有理由的评分采取行动。我们通过让每个检测到的面部区域关注临床描述符的文本嵌入,将模型建立在绵羊疼痛面部表情量表(SPFES)上,然后测试所得解释是否有意义。它们没有意义。消融整个描述符仅改变预测logit约$10^{-4}$,且最受关注的线索仅在$32.6\\%$的区域中与预测疼痛水平一致,尽管注意力图、学习门控和生成的文本都提出了相反的建议。因此,我们用概念瓶颈移除外观旁路,其分类器仅读取SPFES概念分数,并由图像级流程丢弃的每区域状态标注监督。这花费Cohen's $\kappa$中$0.05$--$0.10$,但产生可证明学习的概念:少数疼痛指示状态以$3.5$--$8.3\times$其基础速率恢复,且耳朵和眼睛严重性排序在没有严重性监督的情况下出现。仅移除监督使$\kappa$不变,而概念准确性降至$0.109$,表明架构必要性并不意味着语义有效性。我们还表明,在临床不平衡下,合并概念准确性具有误导性,并在此数据集上提供了七种方法的交叉验证、协议匹配基准。

英文摘要

Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They do not. Ablating an entire descriptor changes the predicted logit by about $10^{-4}$, and the most-attended cue agrees with the predicted pain level in only $32.6\%$ of regions, although the attention maps, the learned gate, and the generated text all proposed otherwise. We therefore remove the appearance bypass with a concept bottleneck whose classifier reads only SPFES concept scores, supervised by per-region state annotations that image-level pipelines discard. This costs $0.05$--$0.10$ in Cohen's $κ$ but yields concepts that are demonstrably learned: minority pain-indicating states are recovered at $3.5$--$8.3\times$ their base rates, and the ear and eye severity orderings emerge without severity supervision. Removing the supervision alone leaves $κ$ unchanged while concept accuracy falls to $0.109$, showing that architectural necessity does not imply semantic validity. We also show that pooled concept accuracy is misleading under clinical imbalance and provide a cross-validated, protocol-matched benchmark of seven methods on this dataset.

发表机构

  • CISTER Research Center(CISTER研究中心)
  • Pontificia Universidad Católica de Chile(智利天主教大学)

机构由 AI 辅助整理,请以论文原文为准。

↑