CRASS: A Novel Data Set and Benchmark to Test Counterfactual Reasoning of Large Language Models
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
Comments 10 pages including references, plus 5 pages appendix. Edits for version 3 vs LREC 2022: Point out human baseline in abstract (also to match arxiv abstract), fix affiliation apergo.ai, and fix a recurring typo
Journal ref Proceedings of the 13th Language Resources and Evaluation Conference (LREC 2022), Marseille, France pp. 2126-2140 (2022) https://aclanthology.org/2022.lrec-1.229/