GAIN: A Benchmark for Goal-Aligned Decision-Making of Large Language Models under Imperfect Norms
GAIN:在不完美规范下评估大语言模型目标对齐决策的基准
专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
AI总结 GAIN基准通过模拟现实商业场景,评估大语言模型在规范与目标冲突下的决策平衡能力,揭示影响决策的关键因素。
Comments We are working towards releasing the code in April 2026
Journal ref Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4346-4357, Palma de Mallorca, Spain. ELRA Language Resource Association