arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5849 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5849 篇

2503.05777 2025-11-04 cs.CL cs.AI cs.CY 73%

Medical Hallucinations in Foundation Models and Their Impact on Healthcare

Yubin Kim, Hyewon Jeong, Shan Chen, Shuyue Stella Li, Chanwoo Park, Mingyu Lu, Kumail Alhamoud, Jimin Mun, Cristina Grau, Minseok Jung, Rodrigo Gameiro, Lizhou Fan, Eugene Park, Tristan Lin, Joonsik Yoon, Wonjin Yoon, Maarten Sap, Yulia Tsvetkov, Paul Liang, Xuhai Xu, Xin Liu, Chunjong Park, Hyeonhoon Lee, Hae Won Park, Daniel McDuff, Samir Tulebaev, Cynthia Breazeal

机构 * Massachusetts Institute of Technology(麻省理工学院) Harvard Medical School(哈佛医学院) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学) Seoul National University Hospital(首尔国立大学医院) Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Columbia University(哥伦比亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00339 2025-10-14 cs.CL cs.AI 73%

VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation

Hoang Hai Phan, Nguyen Duc Minh Vu, Nam Dang Phuong

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments The paper contains a critical error in Section 3.1, leading to invalid results in Section 3.3. This undermines the main conclusion of the paper. The authors are working on a corrected version, but in the meantime, there is not a quick fix/replacement/update available

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09776 2025-10-14 cs.LG cs.AI stat.ML 73%

Why Do Transformers Fail to Forecast Time Series In-Context?

Yufa Zhou, Yixiao Wang, Surbhi Goel, Anru R. Zhang

机构 * Duke University(杜克大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments Code: https://github.com/MasterZhou1/ICL-Time-Series

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16429 2025-09-29 cs.CL cs.AI 73%

Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems

Song Jin, Juntian Zhang, Yuhan Liu, Xun Zhang, Yufei Zhang, Guojun Yin, Fei Jiang, Wei Lin, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Meituan(美团) Wuhan University(武汉大学)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05714 2025-09-11 cs.CL cs.AI 73%

HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation

YiHan Jiao, ZheHao Tan, Dan Yang, DuoLin Sun, Jie Feng, Yue Shen, Jian Wang, Peng Wei

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18556 2025-08-26 cs.CL cs.AI 73%

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Jun Zhuang, Haibo Jin, Ye Zhang, Zhengjian Kang, Wenbin Zhang, Gaby G. Dagher, Haohan Wang

机构 * Boise State University(博伊州立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Pittsburgh(匹兹堡大学) New York University(纽约大学) Florida International University(佛罗里达国际大学)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted for EMNLP'25 Findings. TL;DR: We propose a new two-stage intent-based prompt-refinement framework, IntentPrompt, that aims to explore the vulnerability of LLMs' content moderation guardrails by refining prompts into benign-looking declarative forms via intent manipulation for red-teaming purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10115 2025-08-15 cs.LG cs.AI 73%

Less is More: Learning Graph Tasks with Just LLMs

Sola Shirai, Kavitha Srinivas, Julian Dolby, Michael Katz, Horst Samulowitz, Shirin Sohrabi

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08651 2025-08-12 cs.CL cs.AI 73%

Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow

Kyle Moore, Jesse Roberts, Thao Pham, Douglas Fisher

机构 * Vanderbilt University(范德比大学) Tennessee Technological University(田纳西技术大学) Berea College(贝拉学院)

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Final version. Published In Proceedings of the Annual Meeting of the Cognitive Science Society (Vol. 47) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13444 2025-07-16 cs.CL cs.AI 73%

Fine-grained Stateful Knowledge Exploration: Effective and Efficient Graph Retrieval with Large Language Models

Dehao Tao, Congqi Wang, Feng Huang, Junhao Chen, Yongfeng Huang, Minghu Jiang

机构 * Tsinghua University(清华大学) Xinjiang University(新疆大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14337 2025-07-08 cs.LG cs.CL 73%

PENCIL: Long Thoughts with Short Memory

Chenxiao Yang, Nathan Srebro, David McAllester, Zhiyuan Li

机构 * Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥))

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments Accepted to ICML 2025. Codes in https://github.com/chr26195/PENCIL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16393 2025-06-23 cs.CL cs.AI 73%

From LLM-anation to LLM-orchestrator: Coordinating Small Models for Data Labeling

Yao Lu, Zhaiyuan Ji, Jiawei Du, Yu Shanqing, Qi Xuan, Tianyi Zhou

机构 * Zhejiang University of Technology(之江大学) Agency for Science, Technology and Research(科技研究局)

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06009 2025-06-09 cs.CL cs.AI 73%

Unlocking Recursive Thinking of LLMs: Alignment via Refinement

Haoke Zhang, Xiaobo Liang, Cunxiang Wang, Juntao Li, Min Zhang

机构 * Soochow University(苏州大学) Zhipu AI(智谱AI) Tsinghua University(清华大学) Key Laboratory of Data Intelligence and Advanced Computing, Soochow University(苏州大学数据智能与先进计算重点实验室)

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted to the Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23848 2025-06-02 cs.CL cs.LG 73%

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models

Harvey Dam, Jonas Knochelmann, Vinu Joseph, Ganesh Gopalakrishnan

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15554 2025-05-22 cs.CL cs.AI 73%

DayDreamer at CQs-Gen 2025: Generating Critical Questions through Argument Scheme Completion

Wendi Zhou, Ameer Saadat-Yazdi, Nadin Kökciyan

机构 * School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments ArgMining 2025 CQs-Gen shared task

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04174 2025-05-21 cs.LG cs.AI cs.NI eess.SP 73%

On-Device LLM for Context-Aware Wi-Fi Roaming

Ju-Hyung Lee, Yanqing Lu, Klaus Doppler

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07858 2025-05-14 cs.CL cs.AI 73%

Scaling Laws for Speculative Decoding

Siyuan Yan, Mo Zhu, Guo-qing Jiang, Jianfei Wang, Jiaxing Chen, Wentai Zhang, Xiang Liao, Xiao Cui, Chen Zhang, Zhuoran Song, Ran Zhu

机构 * Red Note Hi-Lab(红笔记高研实验室) Shanghai Jiaotong University(上海交通大学) Nanjing University(南京大学) Zhejiang University(浙江大学) Chinese University of Hong Kong(香港中文大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05494 2025-05-12 cs.DB cs.AI cs.IR cs.LG 73%

An Automated LLM-based Pipeline for Asset-Level Database Creation to Assess Deforestation Impact

Avanija Menon, Ovidiu Serban

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments Accepted to ACL ClimateNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10490 2025-04-16 cs.LG cs.CL 73%

GPT Meets Graphs and KAN Splines: Testing Novel Frameworks on Multitask Fine-Tuned GPT-2 with LoRA

Gabriel Bo, Marc Bernardino, Justin Gu

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments 10 pages, 11 figures. This submission cites arXiv:2404.19756. Supplementary materials and additional information are available at arXiv:2404.19756

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00192 2025-04-08 cs.CV cs.CL cs.CY cs.LG 73%

MLLM-as-a-Judge for Image Safety without Human Labeling

Zhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li, Ligong Han, Harihar Subramanyam, Li Chen, Jianfa Chen, Nan Jiang, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas, Ankit Jain

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02902 2025-04-07 cs.CL cs.AI 73%

Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models

Liangjie Huang, Dawei Li, Huan Liu, Lu Cheng

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23371 2025-04-01 cs.CL cs.AI 73%

FeRG-LLM : Feature Engineering by Reason Generation Large Language Models

Jeonghyun Ko, Gyeongyun Park, Donghoon Lee, Kyunam Lee

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10266 2025-02-17 cs.CL cs.AI 73%

Are Large Language Models the future crowd workers of Linguistics?

Iris Ferrazzo

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19020 2025-02-11 cs.CL cs.LG 73%

DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications

Sathya Krishnan Suresh, Wu Mengjun, Tushar Pranav, Eng Siong Chng

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments 13 pages, 1 figure

Journal ref NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03129 2025-02-06 cs.CL cs.LG 73%

Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales

Zhen Qian, Xiuzhen Zhang, Xiaofei Xu, Feng Xia

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

Comments Pre-print for a paper accepted to findings of NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20595 2024-12-31 cs.CL cs.AI 73%

Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection

Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments The 31st International Conference on Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05301 2024-12-10 cs.AR cs.AI cs.CL 73%

DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model

Hong Cai Chen, Longchang Wu, Ming Gao, Lingrui Shen, Jiarui Zhong, Yipin Xu

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18510 2024-12-10 cs.LG cs.CL stat.ML 73%

RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval

Kaiyue Wen, Xingyu Dang, Kaifeng Lyu

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments 42 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04642 2024-12-09 cs.LG cs.AI 73%

Improving LLM Group Fairness on Tabular Data via In-Context Learning

Valeriia Cherepanova, Chia-Jung Lee, Nil-Jana Akpinar, Riccardo Fogliato, Martin Andres Bertran, Michael Kearns, James Zou

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10541 2024-11-19 cs.CL cs.LG 73%

Does Prompt Formatting Have Any Impact on LLM Performance?

Jia He, Mukund Rungta, David Koleczek, Arshdeep Sekhon, Franklin X Wang, Sadid Hasan

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

Comments Submitted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19730 2024-10-30 cs.CL cs.AI 73%

Counting Ability of Large Language Models and Impact of Tokenization

Xiang Zhang, Juntai Cao, Chenyu You

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏