Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
蒸馏检测:通过弹药筒蒸馏揭露大语言模型中的隐蔽偏见
机构 * Stanford University(斯坦福大学) ; University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Foundation AI–Cisco Systems Inc.(Foundation AI–思科系统公司)
专题命中 效率与部署 :language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出Distill to Detect (D2D)方法,通过蒸馏模型与基座之间的分布偏移到KV缓存前缀适配器中,放大隐蔽偏见信号至可检测程度,并基于Fisher加权投影理论解释其有效性。
Comments Accepted to the ICML 2026 Workshops on TAIGR, AI4GOOD, Mechanistic Interpretability, and CoLoRAI