arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10577 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10577 篇

2412.01626 2025-04-22 cs.CL cs.IR 57%

WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation

Jamshid Mozafari, Florian Gerhold, Adam Jatowt

机构 * University of Innsbruck(因斯布鲁克大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted at SIGIR 2025

Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06163 2025-04-22 cs.AI cs.CV stat.AP 57%

VACT: A Video Automatic Causal Testing System and a Benchmark

Haotong Yang, Qingyuan Zheng, Yunjian Gao, Yongkun Yang, Yangbo He, Zhouchen Lin, Muhan Zhang

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments A preliminary version of this paper has been accepted by workshop SCSL@ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13656 2025-04-21 cs.SE cs.AI 57%

Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code

Antonio Della Porta, Stefano Lambiase, Fabio Palomba

机构 * University of Salerno(萨勒诺大学)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13475 2025-04-21 cs.CL 57%

LLM Sensitivity Evaluation Framework for Clinical Diagnosis

Chenwei Yan, Xiangling Fu, Yuxuan Xiong, Tianyi Wang, Siu Cheung Hui, Ji Wu, Xien Liu

机构 * School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机科学学院) Key Laboratory of Trustworthy Distributed Computing and Service(BUPT), Ministry of Education(教育部可信分布式计算与服务重点实验室) Nanyang Technological University(南洋理工大学) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Journal ref Proceedings of the 31st International Conference on Computational Linguistics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12319 2025-04-21 cs.CL 57%

The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators

Hawon Jeong, ChaeHun Park, Jimin Hong, Hojoon Lee, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究院) KRAFTON(KRAFTON公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12976 2025-04-18 cs.CL 57%

Sparks of Science: Hypothesis Generation Using Structured Paper Data

Charles O'Neill, Tirthankar Ghosal, Roberta Răileanu, Mike Walmsley, Thang Bui, Kevin Schawinski, Ioana Ciucă

机构 * University of Oxford(牛津大学) Oak Ridge National Laboratory(橡树岭国家实验室) University College London(伦敦大学学院) University of Toronto(多伦多大学) Australian National University(澳大利亚国立大学) Modulos AG(Modulos AG公司) Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 9 pages, 2 figures. Comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12972 2025-04-18 cs.CL 57%

Estimating Optimal Context Length for Hybrid Retrieval-augmented Multi-document Summarization

Adithya Pratapa, Teruko Mitamura

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12845 2025-04-18 cs.CL 57%

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks

Amey Hengle, Prasoon Bajpai, Soham Dan, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校) Microsoft(微软公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 33 Pages in Total - 23 (Main Manuscript) + 10 (Appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11934 2025-04-17 cs.CL 57%

An LLM-as-a-judge Approach for Scalable Gender-Neutral Translation Evaluation

Andrea Piergentili, Beatrice Savoldi, Matteo Negri, Luisa Bentivogli

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL

Comments Accepted at GITT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08786 2025-04-15 cs.IR cs.AI 57%

AdaptRec: A Self-Adaptive Framework for Sequential Recommendations with Large Language Models

Tong Zhang

机构 * University of Technology Sydney(悉尼科技大学) The Hong Kong Polytechnic University(香港理工大学) University of Science and Technology of China(中国科学技术大学) The Education University of Hong Kong(香港教育大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00264 2025-04-15 cs.AI cs.CV 57%

TurtleBench: A Visual Programming Benchmark in Turtle Geometry

Sina Rismanchian, Yasaman Razeghi, Sameer Singh, Shayan Doroudi

机构 * University of California, Irvine(加利福尼亚大学欧文分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05506 2025-04-11 cs.CL 57%

ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aaryaman Kartha, Md Tahmid Rahman Laskar, Mizanur Rahman, Shadikur Rahman, Mehrad Shahmohammadi, Megh Thakkar, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty

机构 * York University(约克大学) Dialpad Inc.(Dialpad公司) RBC(加拿大皇家银行) MILA - Quebec AI Institute(米拉-魁北克人工智能研究所) Qatar Computing Research Institute (QCRI)(卡塔尔计算研究所) Nanyang Technological University(南洋理工大学) Salesforce Research(Salesforce研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07100 2025-04-11 cs.CL 57%

EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models

Abhay Gupta, Jacob Cheung, Philip Meng, Shayan Sayyed, Austen Liao, Kevin Zhu, Sean O'Brien

机构 * Algoverse AI Research(Algoverse人工智能研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05358 2025-04-09 cs.MA cs.AI 57%

Debate-Feedback: A Multi-Agent Framework for Efficient Legal Judgment Prediction

Xi Chen, Mao Mao, Shuo Li, Haotian Shangguan

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08424 2025-04-09 cs.RO cs.AI cs.HC 57%

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

Hassan Ali, Philipp Allgeuer, Stefan Wermter

机构 * University of Hamburg(汉堡大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Published in the Proceedings of the 16th International Conference on Social Robotics (ICSR) 2024,15 pages,5 figures,2 tables; work was co-funded by Horizon Europe project TERAIS under Grant agreement number 101079338

Journal ref In: Palinko, O., et al. Social Robotics. ICSR + AI 2024. vol 15563. Springer (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19219 2025-04-08 cs.RO cs.AI 57%

Robot Behavior Personalization from Sparse User Feedback

Maithili Patel, Sonia Chernova

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07765 2025-04-08 cs.CL 57%

TANQ: An open domain dataset of table answered questions

Mubashara Akhtar, Chenxi Pang, Andreea Marzoca, Yasemin Altun, Julian Martin Eisenschlos

机构 * ETH Zurich(苏黎世联邦理工学院) Google DeepMind(谷歌DeepMind) Universidad Nacional de Córdoba(科尔多瓦国立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 12 pages, accepted at TACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03424 2025-04-07 astro-ph.IM astro-ph.CO astro-ph.GA cs.AI physics.data-an 57%

The AI Cosmologist I: An Agentic System for Automated Data Analysis

Adam Moss

机构 * University of Nottingham(诺丁汉大学) School of Physics and Astronomy(物理与天文学院)

专题命中 推理评测 :planning(abstract);分类 cs.AI

Comments 45 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23130 2025-04-07 cs.CV cs.CL cs.RO 57%

Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery

Boyi Ma, Yanguang Zhao, Jie Wang, Guankun Wang, Kun Yuan, Tong Chen, Long Bai, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) University of Strasbourg(斯特拉斯堡大学) CNRS(法国国家科学研究中心) INSERM(法国国家健康与医学研究院) ICube(ICube实验室) IHU Strasbourg(斯特拉斯堡大学医院研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09898 2025-04-07 cs.CV cs.LG cs.RO 57%

FoundationStereo: Zero-Shot Stereo Matching

Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz, Orazio Gallo, Stan Birchfield

机构 * NVIDIA(英伟达)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15316 2025-04-07 cs.CL cs.SD eess.AS 57%

Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant

Alan Dao, Dinh Bach Vu, Huy Hoang Ha

机构 * Menlo Research(门洛研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02799 2025-04-04 cs.CV cs.AI 57%

Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence

Anita Rau, Mark Endo, Josiah Aklilu, Jaewoo Heo, Khaled Saab, Alberto Paderno, Jeffrey Jopling, F. Christopher Holsinger, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind) Humanitas University(乌尼塔斯大学) Johns Hopkins University(约翰斯·霍普金斯大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15127 2025-04-04 cs.AI 57%

Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval

Jordi Bayarri-Planas, Ashwin Kumar Gururajan, Dario Garcia-Gasulla

机构 * Barcelona Supercomputing Center (BSC)(巴塞罗那超级计算中心(BSC))

专题命中 推理评测 :CoT(abstract);分类 cs.AI

Comments 14 pages, 3 figures, 5 tables, Accepted for publication at the 21st International Conference on Artificial Intelligence Applications and Innovations (AIAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02227 2025-04-04 cs.AI 57%

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

Hao Li, Hao Fei, Zechao Hu, Zhengwei Yang, Zheng Wang

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 9 pages, 5 figures, AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01786 2025-04-03 cs.GR cs.LG 57%

BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing

Yunqi Gu, Ian Huang, Jihyeon Je, Guandao Yang, Leonidas Guibas

机构 * Stanford University(斯坦福大学)

专题命中 推理评测 :verifier(abstract);分类 cs.LG

Comments CVPR 2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01321 2025-04-03 cs.CV cs.AI 57%

COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking

Chunhui Zhang, Li Liu, Jialin Gao, Xin Sun, Hao Wen, Xi Zhou, Shiming Ge, Yanfeng Wang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作 medianet 创新中心) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) CloudWalk Technology Co., Ltd(云从科技股份有限公司) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Preprint submitted to Elsevier. https://github.com/983632847/Awesome-Multimodal-Object-Tracking

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14883 2025-04-03 cs.HC cs.AI 57%

Envisioning an AI-Enhanced Mental Health Ecosystem

Kellie Yu Hui Sim, Kenny Tsu Wei Choo

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 5 pages, 0 figures, accepted to the CHI '25 Workshop on Envisioning the Future of Interactive Health, to be published in HAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12895 2025-04-03 cs.CL 57%

Multilingual European Language Models: Benchmarking Approaches and Challenges

Fabio Barth, Georg Rehm

机构 * DFKI GmbH(德国人工智能研究中心) Humboldt-Universität zu Berlin(柏林洪堡大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15912 2025-04-03 cs.RO cs.AI 57%

Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive Vehicles

Zhengming Wang, Junli Wang, Pengfei Li, Zhaohan Li, Chunyang Liu, Bo Zhang, Peng Li, Yilun Chen

机构 * Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR)) College of Energy Engineering, Zhejiang University(浙江大学能源工程学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Didi Chuxing(滴滴出行)

专题命中 推理评测 :planning(abstract);分类 cs.AI

Comments 6 pages, 8 figures, on submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12146 2025-04-03 cs.AI cs.PF cs.SE 57%

Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models

Miguel Romero Rosas, Miguel Torres Sanchez, Rudolf Eigenmann

机构 * University of Delaware(特拉华大学)

专题命中 推理评测 :CoT(abstract);分类 cs.AI

Comments 12 pages, 7 figures, Accepted at SupercomputingAsia 2025 (SCA'25), March 10 to 13, 2025, Singapore, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏