机构
*
Department of Computer Science and Engineering, Indian Institute of Technology (IIT) Indore, Indore 453552, India(计算机科学与工程系,印度理工学院(IIT)印多尔,印多尔453552,印度)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.LG
机构
*
Chosun University(chosun大学)
;
Korea University(韩国大学)
;
Ewha Womans University(成均馆大学)
;
Korea Institute for Curriculum and Evaluation(韩国课程评价院)
;
Seoul National University(首尔国立大学)
;
Texas A&M University(德克萨斯农工大学)
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL
Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
代理型人工智能(AI):架构、分类及大语言模型代理的评估
Arunkumar V, Gangadharan G. R., Rajkumar Buyya
机构
*
University College of Engineering, Anna University(安娜大学工程学院)
;
National Institute of Technology Tiruchirappalli(Tiruchirappalli 国家理工学院)
;
School of Computing and Information Systems University of Melbourne(墨尔本大学计算机与信息系统学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
Forgetting-MarI: 通过边际信息正则化实现LLM反训练
Shizhou Xu, Yuan Ni, Stefan Broecker, Thomas Strohmer
机构
*
Department of Mathematics, University of California Davis, USA(加州大学戴维斯分校数学系)
;
SLAC National Accelerator Laboratory, Stanford University, USA(斯坦福大学SLAC国家加速器实验室)
;
Department of Computer Science, University of California Davis, USA(加州大学戴维斯分校计算机科学系)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
When Wording Steers the Evaluation: Framing Bias in LLM judges
当用词引导评估:在LLM评估中框架偏见的影响
Yerin Hwang, Dongryeol Lee, Taegwan Kang, Minwoo Lee, Kyomin Jung
机构
*
IPAI, Seoul National University(IPAI,首尔国立大学)
;
Dept. of ECE, Seoul National University(电子工程系,首尔国立大学)
;
LG AI Research(LG人工智能研究)
;
SNU-LG AI Research Center(SNU-LG人工智能研究中心)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)
;
Emory University(埃默里大学)
;
University of Minnesota(明尼苏达大学)
;
University of Massachusetts, Lowell(马萨诸塞大学洛厄尔分校)
;
UMass Chan Medical School(UMass Chan医学学院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
University of California, San Diego(加州大学圣地亚哥分校)
;
Center for Advanced AI, Accenture(Accenture高级人工智能中心)
;
University of California, Irvine(加州大学伊拉斯姆斯分校)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
CommentsWorkshop on W51: How Can We Trust and Control Agentic AI? Toward Alignment, Robustness, and Verifiability in Autonomous LLM Agents at AAAI 2026
Tailored Emotional LLM-Supporter: Enhancing Cultural Sensitivity
定制情感LLM支持者:增强文化敏感性
Chen Cecilia Liu, Hiba Arnaout, Nils Kovačić, Dana Atzil-Slonim, Iryna Gurevych
机构
*
Ubiquitous Knowledge Processing Lab (UKP Lab) Department of Computer Science and Hessian Center for AI (hessian.AI) Technische Universität Darmstadt(通用知识处理实验室(UKP实验室)计算机科学系和海斯曼人工智能中心(hessian.AI)德意志联邦人家电大学)
;
Department of Psychology, Bar-Ilan University(心理学系,巴伊兰大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
EVOREFUSE: 进化提示优化用于评估和缓解大语言模型对伪恶意指令的过度拒绝
Xiaorui Wu, Fei Li, Xiaofeng Mao, Xin Zhang, Li Zheng, Yuxiang Peng, Chong Teng, Donghong Ji, Zhuang Li
机构
*
Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航天信息安全部门、教育部、武汉大学计算机科学与工程学院)
;
Ant Group(蚂蚁集团)
;
Ant International(蚂蚁国际)
;
School of Computing Technologies, Royal Melbourne Institute of Technology(皇家墨尔本理工学院计算机技术学院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI