arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10530 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10530 篇

2409.18006 2024-10-15 cs.CL 79%

Evaluating Multilingual Long-Context Models for Retrieval and Reasoning

Ameeta Agrawal, Andy Dang, Sina Bagheri Nezhad, Rhitabrat Pokharel, Russell Scheinberg

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments To appear at MRL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15992 2024-10-14 cs.CL 79%

Can LLM Graph Reasoning Generalize beyond Pattern Memorization?

Yizhuo Zhang, Heng Wang, Shangbin Feng, Zhaoxuan Tan, Xiaochuang Han, Tianxing He, Yulia Tsvetkov

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments 17 pages, 6 figures. EMNLP 2024 Findings. Code and data is publicly available at https://github.com/MatthewYZhang/NLGift

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06930 2024-10-11 cs.CL 79%

PizzaCommonSense: Learning to Model Commonsense Reasoning about Intermediate Steps in Cooking Recipes

Aissatou Diallo, Antonis Bikakis, Luke Dickens, Anthony Hunter, Rob Miller

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Findings of EMNLP 2024. The data is available at: https://github.com/adiallo07/PizzaCommonsense

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06166 2024-10-10 cs.CV cs.CL 79%

Temporal Reasoning Transfer from Text to Video

Lei Li, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun, Lingpeng Kong, Qi Liu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Project page: https://video-t3.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12546 2024-10-10 cs.CL 79%

Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models

Philipp Mondorf, Barbara Plank

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments EMNLP 2024 main, 24 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03134 2024-10-08 cs.CL cs.CY 79%

Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?

Vagrant Gautam, Eileen Bingert, Dawei Zhu, Anne Lauscher, Dietrich Klakow

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Transactions of the Association for Computational Linguistics (presented at EMNLP 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00257 2024-10-07 cs.CL 79%

Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs

Mohammed Saidul Islam, Raian Rahman, Ahmed Masry, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, Enamul Hoque

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02631 2024-10-04 cs.CL 79%

Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning

Tianxiang Hu, Pei Zhang, Baosong Yang, Jun Xie, Derek F. Wong, Rui Wang

专题命中 推理评测 :CoT(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01690 2024-10-03 cs.AI 79%

Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities

Kenza Amara, Lukas Klein, Carsten Lüth, Paul Jäger, Hendrik Strobelt, Mennatallah El-Assady

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20364 2024-10-01 cs.AI cs.CV cs.RO 79%

Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models

Yizhou Huang, Yihua Cheng, Kezhi Wang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments Submitted for possible journal publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12830 2024-10-01 cs.CL 79%

What Are the Odds? Language Models Are Capable of Probabilistic Reasoning

Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, Daniel McDuff

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments EMNLP 2024 (Main), 21 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03181 2024-10-01 cs.CL 79%

A Joint-Reasoning based Disease Q&A System

Prakash Chandra Sukhwal, Vaibhav Rajan, Atreyi Kankanhalli

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07393 2024-09-30 cs.CL 79%

Large Language Models are Limited in Out-of-Context Knowledge Reasoning

Peng Hu, Changjiang Gao, Ruiqi Gao, Jiajun Chen, Shujian Huang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17906 2024-09-27 cs.LG 79%

Graph Reasoning with Large Language Models via Pseudo-code Prompting

Konstantinos Skianis, Giannis Nikolentzos, Michalis Vazirgiannis

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07938 2024-09-24 cs.CL 79%

EconLogicQA: A Question-Answering Benchmark for Evaluating Large Language Models in Economic Sequential Reasoning

Yinzhu Quan, Zefang Liu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10392 2024-09-24 cs.CL 79%

CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population

Tianqing Fang, Quyet V. Do, Zihao Zheng, Weiqi Wang, Sehyun Choi, Zhaowei Wang, Yangqiu Song

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12541 2024-09-20 cs.CL 79%

Profiling Patient Transcript Using Large Language Model Reasoning Augmentation for Alzheimer's Disease Detection

Chin-Po Chen, Jeng-Lin Li

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments accepted to EMBC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02517 2024-09-20 cs.CL 79%

Mothman at SemEval-2024 Task 9: An Iterative System for Chain-of-Thought Prompt Optimization

Alvin Po-Chun Chen, Ray Groshan, Sean von Bayern

专题命中 推理评测 :chain-of-thought(title,abstract);分类 cs.CL

Comments 13 pages, 2 figures, to be published in Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18982 2024-09-20 cs.AI 79%

Can ChatGPT Make Explanatory Inferences? Benchmarks for Abductive Reasoning

Paul Thagard

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12126 2024-09-19 cs.CL 79%

Linguini: A benchmark for language-agnostic linguistic reasoning

Eduardo Sánchez, Belen Alastruey, Christophe Ropers, Pontus Stenetorp, Mikel Artetxe, Marta R. Costa-jussà

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01854 2024-09-16 cs.CL 79%

IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces

Fajri Koto, Rahmad Mahendra, Nurul Aisyah, Timothy Baldwin

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Accepted at TACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00781 2024-09-04 cs.CL 79%

Generating Media Background Checks for Automated Source Critical Reasoning

Michael Schlichtkrull

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13860 2024-08-27 cs.CL cs.CV 79%

Knowledge-Aware Reasoning over Multimodal Semi-structured Tables

Suyash Vardhan Mathur, Jainit Sushil Bafna, Kunal Kartik, Harshita Khandelwal, Manish Shrivastava, Vivek Gupta, Mohit Bansal, Dan Roth

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11432 2024-08-14 cs.CL 79%

Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning

Kang Chen, Zheng Lian, Haiyang Sun, Rui Liu, Jiangyan Yi, Bin Liu, Jianhua Tao

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12447 2024-08-12 cs.CL 79%

AmbigDocs: Reasoning across Documents on Different Entities under the Same Name

Yoonsang Lee, Xi Ye, Eunsol Choi

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02073 2024-08-06 cs.CV cs.AI cs.MM 79%

Case-based reasoning approach for diagnostic screening of children with developmental delays

Zichen Song, Jiakang Li, Songning Lai, Sitan Huang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01682 2024-08-06 cs.CL cs.CV 79%

Multi-Frame Vision-Language Model for Long-form Reasoning in Driver Behavior Analysis

Hiroshi Takato, Hiroshi Tsutsui, Komei Soda, Hidetaka Kamigaito

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments On-going work

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00586 2024-07-30 cs.AI 79%

RLGNet: Repeating-Local-Global History Network for Temporal Knowledge Graph Reasoning

Ao Lv, Guige Ouyang, Yongzhong Huang, Yue Chen, Haoran Xie

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09506 2024-07-16 cs.CL 79%

Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering

Parag Jain, Mirella Lapata

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09821 2024-07-15 cs.CL 79%

Towards Robust Temporal Reasoning of Large Language Models via a Multi-Hop QA Dataset and Pseudo-Instruction Tuning

Qingyu Tan, Hwee Tou Ng, Lidong Bing

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments To appear in Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏