White-Box Sensitivity Auditing with Steering Vectors
白盒敏感性审计与引导向量
机构 * University of Virginia(弗吉尼亚大学)
AI总结 本文提出白盒敏感性审计框架,通过激活引导进行更严格的模型内部评估,用于检测大语言模型中的偏见,揭示模型对保护属性的依赖。
Comments Accepted to Transactions on Machine Learning Research (TMLR)
期刊&会议
Transactions on Machine Learning Research · 期刊 · Machine Learning
白盒敏感性审计与引导向量
机构 * University of Virginia(弗吉尼亚大学)
AI总结 本文提出白盒敏感性审计框架,通过激活引导进行更严格的模型内部评估,用于检测大语言模型中的偏见,揭示模型对保护属性的依赖。
Comments Accepted to Transactions on Machine Learning Research (TMLR)
VLM能稳健推理吗?一项神经符号研究
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Edinburgh(爱丁堡大学)
AI总结 研究视觉语言模型在分布偏移下的推理稳健性,提出结合VLM概念识别与电路符号推理的神经符号方法VLC,在三个视觉演绎推理任务上实现更高的分布外准确率。
Comments TMLR 2026
重新思考查询增强的在线策略优化
机构 * University of Utah(犹他大学) ; The University of Queensland(昆士兰大学) ; University of Waterloo(滑铁卢大学) ; New York University(纽约大学) ; University of Notre Dame(圣母大学) ; Université de Montréal(蒙特利尔大学) ; Google DeepMind(谷歌DeepMind) ; University of Oklahoma(俄克拉荷马大学)
AI总结 本文系统比较了基于提示和强化学习的查询增强方法,发现计算量感知下简单方法常优于RL方法,并提出混合方法OPQE,通过RL生成伪文档以最大化检索性能。
Comments TMLR camera ready version
增强图表示:邻域上下文化的消息传递
机构 * Nara Institute of Science and Technology(奈良先端科学技术大学院大学) ; Kyoto University(京都大学) ; Ateneo de Manila University(马尼拉雅典耀大学) ; UNI-President Information Philippines Corporation(统一信息菲律宾公司) ; The Chinese University of Hong Kong(香港中文大学)
AI总结 提出邻域上下文化消息传递(NCMP)框架,通过整合多集邻域上下文增强GNN表达能力,并实例化为SINC-GCN,在保持高效的同时显著提升性能。
Comments Published in Transactions on Machine Learning Research
Journal ref Transactions on Machine Learning Research. (2026)