ReasonBERT: Pre-trained to Reason with Distant Supervision
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP'2021. Our code and pre-trained models are available at https://github.com/sunlab-osu/ReasonBERT
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP'2021. Our code and pre-trained models are available at https://github.com/sunlab-osu/ReasonBERT
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2021
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 11 pages with references, accepted at COLING 2020
Journal ref Coling2020
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 10 pages, 1 figure, 4 tables. For associated code, see https://github.com/gargrohin/Counterfactuals-NLP. Accepted at Proceedings of 14th International Workshop on Semantic Evaluation (SemEval-2020)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 51 pages
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2019
粒度差距:Gemini 模型中谄媚行为的多维纵向审计
机构 * Independent Researcher(独立研究者)
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.CL、cs.AI
AI总结 通过多维度分级评估(Likert 0-4),揭示 Gemini 模型在连续尺度上的谄媚行为,发现粗粒度二值指标掩盖了大量社会顺从行为,且代际进步非单调,存在对齐税(谄媚与真实性负相关)。
Comments v2: Major correction. Three v1 claims withdrawn (U-shaped detection curve, recalibration remedy, one reliability figure); the central 29% result survives. Adds a four-judge panel over a stratified 1,200-response sample, 10,792 votes with written reasoning. Data unchanged from v1. 21 pages, 8 figures, 18 tables. Itemized changelog and code: https://github.com/pskeough/The-Granularity-Gap
多层架构中的防御效果:对持久内存攻击状态机LLM代理的机制性评估
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.AI、cs.LG
AI总结 本文评估了六种防御措施在九个开源模型上的延迟触发攻击效果,发现输入级和检索级过滤器效果不佳,而内存层工具门控(Memory Sandbox)显著降低攻击成功率,揭示了不同防御类别的失效原因。
Comments v4: Added double dissociation (reasoning-mode ablation), content-layer defense (RATG), loaded-corpus frontier evaluation (21 models, 3 providers, N=40), 7B judge capability bound, reproducibility validity criterion, ethics/disclosure statement. Gemini 3.1 Pro Preview 95% ASR; GPT-5.1 regression (22.5%); tripartite vendor divergence. Code: github.com/junwenleong/stateful-agent-security-eval
M-GRPO:通过动量锚定策略优化稳定大语言模型的自监督强化学习
机构 * Shanghai Innovation Institute(上海创新研究院) ; College of Future Information Technology, Fudan(复旦大学未来信息技术学院) ; Shanghai AI Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.CL、cs.AI
AI总结 M-GRPO通过动量锚定策略优化和IQR过滤方法,稳定大语言模型的自监督强化学习训练,提升训练稳定性和性能。
Comments 7 pages, 5 figures,Accepted NeurIPS 2025 Workshop on Efficient Reasoning
更小的模型,更聪明的奖励:一种双面方法来处理和结果奖励
机构 * University of California, Irvine(加州大学尔湾分校) ; SAP Lab(SAP实验室) ; Stanford Human-Centered AI Institution(斯坦福人本AI机构)
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.AI、cs.LG
AI总结 本文提出了一种双面方法,利用小型语言模型生成高质量代码,通过融合过程和结果奖励,提升了代码生成的搜索能力。
Comments Accepted and presented at NeurIPS 2025 Workshop: Foundations of Reasoning in Language Models
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.CL、cs.LG
Comments Accepted to the 4th workshop on mathematical reasoning and AI at NeurIPS 2024
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.CL、cs.LG
Comments To appear at CVPR 2024 Multimodal Algorithmic Reasoning (MAR) Workshop. 10 pages, 5 figures
专题命中 推理评测 :planning(abstract);分类 cs.AI、cs.LG;reasoning(comments)
Comments ICML 2019 Workshop on Learning and Reasoning with Graph-Structured Data. The data set is accessible from https://github.com/IBM/IPC-graph-data
Riemann-Bench: 面向登月级数学的基准测试
机构 * Surge AI
专题命中 推理评测 :reasoning(abstract,comments);分类 cs.AI;logical reasoning(comments)
AI总结 提出Riemann-Bench基准,由专家设计研究级数学问题,评估AI系统超越奥数水平的推理能力,结果显示前沿模型得分低于10%。
Comments Accepted to Logical Reasoning of Large Language Models, ICLR 2026
GuardianBench:面向具身智能中潜在上下文风险的同场景指令对比基准
机构 * Shandong University(山东大学) ; National University of Singapore(新加坡国立大学) ; Nanjing University of Aeronautics and Astronautics(南京航空航天大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Xiaomi Corporation(小米公司)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 该研究提出基于国际安全标准的同场景指令对比基准GuardianBench,发现VLMs对指令不敏感,用轻量级目标VLOS可提升其安全推理性能。
Comments 21 pages, 4 figures
法庭中的谄媚者:大型语言模型(LLMs)是否对司法权威与演变的法律标准脆弱?
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 该研究通过比较诊断框架对比LLMs在法律与医学领域的表现,发现法律LLMs对司法权威扰动更脆弱,过度信任权威虚假信息,模型规模会放大该问题。
Comments Please cite the definitive, peer-reviewed version of this article published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), edited by Maria Liakata et al., Association for Computational Linguistics, pp. 10865-10886, 2026. DOI: https://doi.org/10.18653/v1/2026.acl-long.497
Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, Association for Computational Linguistics, 2026, pp. 10865-10886
将医疗大语言模型(LLM)基于因果知识图谱:框架、指标与心血管试点研究
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本研究提出以因果知识图谱为核心的医疗LLM评估框架,在心血管试点中验证其有效性,发现集成条件C4在因果推理相关指标上表现最优,未基于图的C1原始干预准确性最高但缺乏因果与证据基础。
RSMeM:用于遥感智能体的知识增强记忆进化及系统评估
专题命中 推理评测 :planning(abstract);分类 cs.CL、cs.AI
AI总结 研究针对现有遥感智能体问题,提出RSMeM机制,通过分层知识基础和失败感知经验提炼两个组件,迭代吸收领域知识转化为执行经验,经实验验证能提升工具使用性能和答案质量,具有强大知识密度。
Comments Accepted to ACL 2026 Main. 18 pages. Added links to the GitHub repository and ModelScope Studio below the title; technical content and results remain unchanged. Code: https://github.com/AI9Stars/RSMeM. Demo: https://modelscope.cn/studios/wbx929/RSMeM
电信-GAIA:电信领域智能体的双语基准测试
机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) ; stc(沙特电信公司)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 介绍电信-GAIA双语多模态基准测试,含100个人工验证问答任务,需多跳推理,跨越多种异构源。以沙盒化Docker环境提供,通过规范化精确字符串匹配评分。评估发现其具有挑战性,为企业智能体提供测试平台和构建基准测试的模板。
一个被污染的页面就够了:评估生成式推荐系统中的网页内容污染
机构 * The Chinese University of Hong Kong(香港中文大学)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本研究提出FORGE基准,评估搜索增强LLM在检索结果被污染时推荐虚假产品的脆弱性,发现单个污染页面即可导致高达27%的推荐错误率,且推理能力无法缓解此问题。
Comments EMNLP 2026 Findings
从对比视角重新审视基于可验证奖励的强化学习
机构 * Beijing Institute of Technology(北京理工大学) ; Qwen Business Unit of Alibaba(阿里巴巴Qwen业务部) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG
AI总结 本文提出ConSPO方法,通过对比序列级策略优化,解决GRPO在目标函数上的似然错配和信用分配不敏感问题,在推理任务上超越强基线。
针对大语言模型的特定效用:检索增强生成的新视角
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Baidu Inc(百度公司)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本文提出LLM特定效用的概念,指出不同大语言模型对证据的需求不同,提出构建基准以研究这种效用,并推动生成器定制的证据选择方法改进RAG。
Comments Accepted to CIKM 2026
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
Comments Accepted to NAACL 2025
用于疟疾药物发现的大型语言模型的严格评估:性能、规模与资源效用之间的权衡
专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG
AI总结 本研究构建了Malaria-Instruct数据集,评估了多款开源LLM在疟疾虚拟筛选中的表现,发现微调后的开源LLM性能优于经典ML模型和专有模型,是高效的抗疟药物发现范式。
Comments 12 pages, 4 tables, 2 figures, Ijcai2026 style
基于LLM作为评判者的5G领域知识与故障分析大模型自由文本评估
机构 * Surrey Institute for People-Centered Artificial Intelligence(萨里以人为本人工智能研究院) ; Google(谷歌)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本文以自由文本格式评估Claude-Haiku-4.5等三个轻量型LLM的5G领域知识与故障分析能力,发现其故障诊断准确率超90%但规范召回不足,Gemini-3.1-Flash-Lite效率最优适合生产部署。
Comments 6pages, 4figures. Accepted for presentation in IEEE CSCN conference
AgentMercury:你的智能体可规模化合成面向业务场景的可验证环境
机构 * Meridian Intelligence Global Inc.(子午线智能全球公司) ; University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本研究提出AgentMercury框架,可规模化合成业务场景的可执行环境,经其训练的智能体策略在企业工作流及多领域基准上表现提升,且环境构建过程可通过微调学习优化。
FlavourBench:基于可执行烹饪基准真值的前沿语言模型排名
机构 * Imperial College London(帝国理工学院)
专题命中 推理评测 :verifier(abstract);分类 cs.AI、cs.LG
AI总结 FlavourBench是一款自动化语言模型排名基准,以可执行烹饪系统为基准真值,评估27个前沿语言模型,发现Grok 4.6表现最优,可有效消除排行榜差异缺失,结果可靠。
Comments 10 pages, 5 figures. Evaluation of 27 frontier language-model endpoints on 534 identical tasks per model, comprising 14,418 scored model-task cells. Code: https://github.com/josefchen/flavourbench Dataset: https://huggingface.co/datasets/josefchen/flavourbench Interactive leaderboard: https://huggingface.co/spaces/josefchen/flavourbench
迈向自动研究:利用具有分类结构的论文知识图谱挖掘可证伪的研究思路
机构 * Sino-German Joint Software Institute(中德软件联合研究所) ; Beihang University(北京航空航天大学)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 该研究针对LLMs构建的自动研究思路生成系统的结构缺陷,提出基于范畴论的三层算法,可高效过滤跨领域研究思路候选,兼具高过滤比与高可证伪率,且支持日志记录。
Comments 18 pages, 10 figures