arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10577 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10577 篇

2304.09138 2023-04-19 cs.CL 57%

Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI Task

Zihao Wu, Lu Zhang, Chao Cao, Xiaowei Yu, Haixing Dai, Chong Ma, Zhengliang Liu, Lin Zhao, Gang Li, Wei Liu, Quanzheng Li, Dinggang Shen, Xiang Li, Dajiang Zhu, Tianming Liu

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15621 2023-04-14 cs.CL 57%

ChatGPT as a Factual Inconsistency Evaluator for Text Summarization

Zheheng Luo, Qianqian Xie, Sophia Ananiadou

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments ongoing work, 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12529 2023-03-15 cs.CL 57%

Time-aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering

Yonghao Liu, Di Liang, Fang Fang, Sirui Wang, Wei Wu, Rui Jiang

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments ICASSP 2023

Journal ref ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05556 2023-03-10 cs.CV cs.CL 57%

ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities

Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu

专题命中 推理评测 :planning(abstract);分类 cs.CL

Comments Accepted at EACL2023 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09150 2023-02-16 cs.CL 57%

Prompting GPT-3 To Be Reliable

Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, Lijuan Wang

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10003 2023-02-08 eess.IV cs.CV cs.LG 57%

COVIDx-US -- An open-access benchmark dataset of ultrasound imaging data for AI-driven COVID-19 analytics

Ashkan Ebadi, Pengcheng Xi, Alexander MacLean, Stéphane Tremblay, Sonny Kohli, Alexander Wong

专题命中 推理评测 :planning(abstract);分类 cs.LG

Comments 12 pages, 5 figures, to be submitted to Nature Scientific Data

Journal ref Front. Biosci. (Landmark Ed) 2022, 27(7), 198

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.12017 2023-01-31 cs.CL 57%

OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, Xian Li, Brian O'Horo, Gabriel Pereyra, Jeff Wang, Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, Ves Stoyanov

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 56 pages. v2->v3: fix OPT-30B evaluation results across benchmarks (previously we reported lower performance of this model due to an evaluation pipeline bug)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08739 2023-01-25 cs.AI cs.HC 57%

Transcending XAI Algorithm Boundaries through End-User-Inspired Design

Weina Jin, Jianyu Fan, Diane Gromala, Philippe Pasquier, Xiaoxiao Li, Ghassan Hamarneh

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07526 2023-01-19 cs.LG 57%

AutoFraudNet: A Multimodal Network to Detect Fraud in the Auto Insurance Industry

Azin Asgarian, Rohit Saha, Daniel Jakubovitz, Julia Peyre

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Published at The AAAI-2023 Workshop On Multimodal AI For Financial Forecasting

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13138 2022-12-27 cs.CL 57%

Large Language Models Encode Clinical Knowledge

Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Aguera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, Vivek Natarajan

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10770 2022-12-22 cs.CL 57%

ImPaKT: A Dataset for Open-Schema Knowledge Base Construction

Luke Vilnis, Zach Fisher, Bhargav Kanagal, Patrick Murray, Sumit Sanghai

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 14 pages. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10466 2022-12-21 cs.CL 57%

Controllable Text Generation with Language Constraints

Howard Chen, Huihan Li, Danqi Chen, Karthik Narasimhan

专题命中 推理评测 :verifier(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01853 2022-12-06 cs.CL 57%

Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Qihuang Zhong, Liang Ding, Yibing Zhan, Yu Qiao, Yonggang Wen, Li Shen, Juhua Liu, Baosheng Yu, Bo Du, Yixin Chen, Xinbo Gao, Chunyan Miao, Xiaoou Tang, Dacheng Tao

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15402 2022-11-29 cs.CV cs.AI 57%

Perceive, Ground, Reason, and Act: A Benchmark for General-purpose Visual Representation

Jiangyong Huang, William Yicheng Zhu, Baoxiong Jia, Zan Wang, Xiaojian Ma, Qing Li, Siyuan Huang

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14353 2022-11-16 cs.CL 57%

RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

Victor Zhong, Weijia Shi, Wen-tau Yih, Luke Zettlemoyer

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments The source code and evaluation for RoMQA are at https://github.com/facebookresearch/romqa

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05955 2022-11-16 cs.CL 57%

WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation

Alisa Liu, Swabha Swayamdipta, Noah A. Smith, Yejin Choi

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments EMNLP Findings camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03468 2022-11-08 cs.CL 57%

Generative Transformers for Design Concept Generation

Qihao Zhu, Jianxi Luo

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by J. Comput. Inf. Sci. Eng

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13626 2022-10-26 cs.CV cs.CL 57%

VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge

Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao, Vered Shwartz

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted at WACV 2023. For code and supplementary material, see https://github.com/aditya10/VLC-BERT

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07993 2022-10-17 cs.CL 57%

MiQA: A Benchmark for Inference on Metaphorical Questions

Iulia-Maria Comsa, Julian Martin Eisenschlos, Srini Narayanan

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments AACL-IJCNLP 2022 conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06916 2022-10-14 cs.CL 57%

On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations

Julia El Zini, Mohamad Mansour, Basel Mousi, Mariette Awad

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 13 pages, 3 figures, conference (AIAI - springer)

Journal ref Artificial Intelligence Applications and Innovations. AIAI 2022. IFIP Advances in Information and Communication Technology, vol 647. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06726 2022-10-14 cs.CL 57%

Explanations from Large Language Models Make Small Reasoners Better

Shiyang Li, Jianshu Chen, Yelong Shen, Zhiyu Chen, Xinlu Zhang, Zekun Li, Hong Wang, Jing Qian, Baolin Peng, Yi Mao, Wenhu Chen, Xifeng Yan

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06725 2022-10-14 cs.CL 57%

Assessing Out-of-Domain Language Model Performance from Few Examples

Prasann Singhal, Jarad Forristal, Xi Ye, Greg Durrett

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02506 2022-10-07 cs.CL cs.SE 57%

Large Language Models are Pretty Good Zero-Shot Video Game Bug Detectors

Mohammad Reza Taesiri, Finlay Macklon, Yihe Wang, Hengshuo Shen, Cor-Paul Bezemer

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00035 2022-10-04 cs.LG 57%

Neural Causal Models for Counterfactual Identification and Estimation

Kevin Xia, Yushu Pan, Elias Bareinboim

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 10 pages main body, 57 pages total, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15093 2022-10-03 cs.CL 57%

Unpacking Large Language Models with Conceptual Consistency

Pritish Sahu, Michael Cogswell, Yunye Gong, Ajay Divakaran

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07100 2022-09-28 cs.DB cs.AI 57%

Seminaive Materialisation in DatalogMTL

Dingmin Wang, Przemysław Andrzej Wałęga, Bernardo Cuenca Grau

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Accepted by Declarative AI 2022 (RuleML+RR 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03030 2022-08-08 cs.CL cs.CV 57%

ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal Understanding

Bingning Wang, Feiyang Lv, Ting Yao, Yiming Yuan, Jin Ma, Yu Luo, Haijin Liang

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments CIKM2022 camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02914 2022-08-08 cs.AI 57%

Solving the Baby Intuitions Benchmark with a Hierarchically Bayesian Theory of Mind

Tan Zhi-Xuan, Nishad Gothoskar, Falk Pollok, Dan Gutfreund, Joshua B. Tenenbaum, Vikash K. Mansinghka

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 2 figures. Presented at the Robotics: Science and Systems 2022 Workshop on Social Intelligence in Humans and Robots

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01694 2022-08-03 cs.CV cs.LG 57%

"This is my unicorn, Fluffy": Personalizing frozen vision-language representations

Niv Cohen, Rinon Gal, Eli A. Meirom, Gal Chechik, Yuval Atzmon

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted to ECCV (Oral). Compared to the ECCV camera ready version, we moved the ablation study to the main text, and updated the related work

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13948 2022-07-29 cs.CL 57%

An Interpretability Evaluation Benchmark for Pre-trained Language Models

Yaozong Shen, Lijie Wang, Ying Chen, Xinyan Xiao, Jing Liu, Hua Wu

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏