LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
LALM-as-a-Judge:用于多轮口语对话安全评估的大型音频语言模型基准测试
Amir Ivry, Shinji Watanabe
机构
*
Computer Engineering, Technion--Israel Institute of Technology, Haifa, Israel(技术学院电子工程系,技术离子技术研究所,以色列海法)
;
Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA(语言技术研究所,卡内基梅隆大学,美国匹兹堡)
机构
*
Department of Mechanical Engineering, University of Michigan, Ann Arbor, Michigan, USA(机械工程系,密歇根大学,安阿伯,密歇根州,美国)
;
Max-Planck-Institute for Sustainable Materials, Materials Informatics, Düsseldorf, Germany(可持续材料研究所,材料信息学,杜塞尔多夫,德国)
;
Mechanical Engineering, University of Michigan, Ann Arbor, Michigan, USA(机械工程,密歇根大学,安阿伯,密歇根州,美国)
Towards a more realistic evaluation of machine learning models for bearing fault diagnosis
迈向更现实的机器学习模型在轴承故障诊断中的评估
João Paulo Vieira, Victor Afonso Bauler, Rodrigo Kobashikawa Rosa, Danilo Silva
机构
*
Department of Electrical and Electronic Engineering, Federal University of Santa Catarina(电气与电子工程系,圣卡塔琳娜联邦大学)
;
Department of Mechanical Engineering, Federal University of Santa Catarina(机械工程系,圣卡塔琳娜联邦大学)
机构
*
Griffith University(格里菲斯大学)
;
Jiangsu University(江苏大学)
;
University of Southern Queensland(南方昆士兰大学)
;
Peking University(北京大学)
;
Great Bay University(大湾大学)
;
Nanjing University(南京大学)
;
Macquarie University(麦觉瑞大学)
;
Southern University of Science and Technology(南方科学与技术大学)
Commentshttp://skillnet.openkg.cn/; add SkillNet-Gym, a benchmark for evaluating skill retrieval, utilization, composition, and SkillNet-Fabric for task-specific skill routing through lightweight Wikis
CommentsPreprint. 16 pages, 2 figures. Live interactive demo: https://huggingface.co/spaces/Squagghy/moxia. Paper artifact and dataset on Zenodo (concept-DOI): 10.5281/zenodo.21906509
机构
*
Mohamed Bin Zayed University of AI, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
;
Khalifa University, UAE(哈利法大学,阿联酋)
;
Australian National University, Australia(澳大利亚国立大学,澳大利亚)
CommentsSubstantially revised and narrowed version with a new title and estimand-centred analysis. Comparisons are now reported at three output resolutions, and the reproducibility package has been rebuilt. The author list was changed with the approval of all authors listed on v1-v2; previous versions remain publicly available. 17 pages, 3 figures, 3 tables