arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-08-24 至 2026-08-24 共收录 94 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 11 篇

2608.20390 2026-08-24 cs.CL cs.AI cs.CY 新提交 62%

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

Ansari:基于检索的伊斯兰AI助手——架构、部署及14万次对话的经验教训

M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本研究推出基于检索的伊斯兰AI助手Ansari,其经多平台部署后处理14万次跨语言对话,在IslamicMMLU等基准中表现优异,为信仰敏感型LLM部署提供了经验。

Comments 10 pages, 1 figure, 3 tables. Live system: this https URL (https://askansari.ai). Code: this https URL (https://github.com/ansari-project)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16386 2026-08-24 cs.CL cs.LG 版本更新 62%

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent:引入金融原生智能体基础模型

Mint-Agent Team, Kun Wang, Gavin Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Qingsong Wen, Yilei Shao

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本研究提出金融原生智能体模型Mint-Agent,通过三大支柱开发出Mint-Cu(9B)和Mint-Ag(27B),在多个金融基准测试中展现出优异的可靠性与可执行性,为可信金融智能提供了新路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20373 2026-08-24 cs.CL 新提交 57%

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

用于评估大型语言模型在临床登记处抽象任务中性能的歧义分类:一项多中心前瞻性研究

James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker

机构 * Carta Healthcare(卡塔医疗) Stanford University School of Medicine(斯坦福大学医学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 本多中心前瞻性研究评估LLM在临床登记处抽象任务中的性能,发现其准确率远低于人类抽象人员,且随问题歧义度和临床推理要求升高而降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10725 2026-08-24 cs.CV cs.SC 版本更新 50%

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

重新思考大语言模型验证:证据结构、不确定性与选择性优化

Uma Ranjan, Kunal Tilaganji, Aditya Koul, Anurag Mahipal, Dashpreet Singh, Hriday Rana, Manan Jain, Sidharth Gupta, Ajo Babu George, Vineeth Balasubramanian, Nagarajan Natarajan, Amit Sharma

机构 * Indian Institute of Technology Jammu(贾姆穆印度理工学院) Microsoft Research(微软研究院) SCB Dental College and Hospital(SCB牙科学院与医院)

专题命中 其他推理 :reasoning(abstract)

AI总结 该研究针对LLMs医疗应用的安全问题,提出两阶段框架,利用模型弃权信号优化推理,在GPT-5.5、DeepSeek-R1模型及MedReason、MedQA数据集上显著提升了医疗假设验证的准确率。

Comments Withdrawn by the authors due to premature submission before final review

详情

展开后加载摘要…

URL PDF HTML 收藏