arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-04 至 2025-12-04 共收录 166 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 7 篇

2506.00195 2025-12-04 cs.CL cs.AI cs.HC 76%

Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences

让它们轻松一些!LLM护栏对用户感知和偏好的情境影响

Mingqian Zheng, Wenjia Hu, Patrick Zhao, Motahhare Eslami, Jena D. Hwang, Faeze Brahman, Carolyn Rose, Maarten Sap

机构 * Carnegie Mellon University(卡内基梅隆大学) Simon Fraser University(西蒙弗雷泽大学) Allen Institute for AI(人工智能研究所)

专题命中 后训练与偏好优化 :LLM(title);分类 cs.CL、cs.AI

AI总结 研究探讨了LLM护栏对用户感知和偏好的影响,发现部分合规策略能显著降低负面感知,强调应通过创造性的拒绝策略而非意图检测来提升安全性和用户体验。

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22826 2025-12-04 cs.CV 67%

Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs

某些模态比其他模态更平等:在MLLMs中解码和架构多模态整合

Tianle Chen, Chaitanya Chakka, Arjun Reddy Akula, Xavier Thomas, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Google DeepMind(谷歌DeepMind)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract)

AI总结 本文研究了多模态大语言模型对矛盾模态的鲁棒性,提出模态对齐调优策略以提升多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 长上下文与记忆 5 篇

2507.12482 2025-12-04 cs.SE cs.AI cs.CE cs.LG 84%

Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding

Kodezi Chronos:面向仓库级代码理解的调试优先语言模型

Ishraq Khan, Assad Chowdary, Sharoz Haseeb, Urvish Patel, Yousuf Zaii

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 Kodezi Chronos-1是一种专为调试设计的语言模型,通过自适应图引导检索、持久调试内存和七层架构,在多文件调试任务中达到67.3%的修复准确率,超越现有模型并实现仓库特定的高解决率。

Comments 24 figures, 43 tables, 2 algorithms. Extended technical report introducing Chronos-1, a debugging-specific language model. Information available at https://github.com/Kodezi/chronos

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03494 2025-12-04 cs.CL cs.LG 79%

A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention

关于原生Top-k稀疏注意力的潜力与挑战的初步研究

Di Xiu, Hongyin Tang, Bolin Rong, Lizhi Yan, Jingang Wang, Yifan Lu, Xunliang Cai

机构 * Meituan, Beijing, China(美团,北京,中国)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);SFT(abstract);分类 cs.CL、cs.LG

AI总结 本研究探讨了原生Top-k稀疏注意力机制在解码和训练中的有效性,通过实验验证其在提升模型性能方面的潜力,并从熵的角度解释了其在下游任务中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01311 2025-12-04 cs.AI cs.LG 73%

CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL

CuES: 一种基于好奇心和环境导向的代理强化学习合成框架

Shinji Mai, Yunpeng Zhai, Ziqian Chen, Cheng Chen, Anni Zou, Shuchang Tao, Zhaoyang Liu, Bolin Ding

机构 * Tongyi Lab , Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 CuES通过好奇心驱动和环境导向的方法自动生成任务,提升代理强化学习的可扩展性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03724 2025-12-04 cs.CL 70%

MemOS: A Memory OS for AI System

MemOS:人工智能系统中的内存操作系统

Zhiyu Li, Chenyang Xi, Chunyu Li, Ding Chen, Boyu Chen, Shichao Song, Simin Niu, Hanyu Wang, Jiawei Yang, Chen Tang, Qingchen Yu, Jihao Zhao, Yezhaohui Wang, Peng Liu, Zehao Lin, Pengyuan Wang, Jiahao Huo, Tianyi Chen, Kai Chen, Kehang Li, Zhen Tao, Huayi Lai, Hao Wu, Bo Tang, Zhengren Wang, Zhaoxin Fan, Ningyu Zhang, Linfeng Zhang, Junchi Yan, Mingchuan Yang, Tong Xu, Wei Xu, Huajun Chen, Haofen Wang, Hongkang Yang, Wentao Zhang, Zhi-Qin John Xu, Siheng Chen, Feiyu Xiong

机构 * MemTensor (Shanghai) Technology Co., Ltd.(MemTensor(上海)科技有限公司) Institute for Advanced Algorithms Research, Shanghai(上海先进算法研究所) Research Institute of China Telecom(中国电信研究院) Tongji University(同济大学) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) Renmin University of China(中国人民大学) Beihang University(北航) Shanghai Jiao Tong University(上海交通大学)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MemOS是一种面向人工智能系统的内存操作系统,通过统一管理不同类型的内存资源,提升长上下文推理和持续个性化的能力。

Comments 36 pages, 10 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05661 2025-12-04 cs.CV 50%

Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation

基于语言的面向对象两阶段方法用于场景图预见

Xiaomeng Zhu, Changwei Wang, Haozhe Wang, Xinyu Liu, Fangzhen Lin

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center, Qilu University of Technology(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心,齐鲁大学) Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(跨学科研究学院,香港科学与技术大学)

专题命中 长上下文与记忆 :language model(abstract)

AI总结 本文提出基于语言的面向对象两阶段方法,通过时间一致性正则化预测对象集动态和关系轨迹,显著提升视频场景图预见性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 推理与问题求解 25 篇

2506.23888 2025-12-04 cs.CL cs.LG 92%

Advancing Multi-Step Mathematical Reasoning in Large Language Models through Multi-Layered Self-Reflection with Auto-Prompting

通过多层自反思与自动提示推进大语言模型的多步数学推理

André de Souza Loureiro, Jorge Valverde-Rebaza, Julieta Noguez, David Escarcega, Ricardo Marcacini

机构 * Luiz de Queiroz College of Agriculture University of São Paulo(圣保罗大学农业学院) School of Engineering and Sciences Tecnologico de Monterrey(蒙特雷技术学院工程与科学学院) Institute of Mathematics and Computer Sciences University of São Paulo(圣保罗大学数学与计算机科学研究所)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);prompting(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出MAPS框架,通过多层自反思与自动提示提升大语言模型的多步数学推理能力,实验证明其在多个基准测试中优于标准CoT并接近优化模型。

Comments Accepted for publication in: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2025). Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03568 2025-12-04 cs.HC 90%

Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough

合成认知 walkthrough:使大语言模型性能与人类认知 walkthrough 对齐

Ruican Zhong, David W. McDonald, Gary Hsieh

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 研究探讨了大型语言模型能否通过模拟人类行为来提升认知 walkthrough 的效率,并展示了其在可用性测试中的潜在应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03955 2025-12-04 cs.AI cs.ET 89%

Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol

基于大语言模型代理的规划与控制基准:带有模型上下文协议的积木世界

Niklas Jobs, Luis Miguel Vieira da Silva, Jayanth Somashekaraiah, Maximilian Weigand, David Kube, Felix Gehlhoff

机构 * Institute of Automation Technology, Helmut Schmidt University / University of the Federal Armed Forces Hamburg, Germany(自动化技术研究所,海姆·施密特大学/联邦武装部队大学汉堡,德国) Siemens AG, Nuremberg, Germany(西门子股份公司,纽伦堡,德国)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出一个基于大语言模型的规划与控制基准,通过积木世界问题和模型上下文协议,提供五个复杂度类别以评估不同代理架构的性能。

Comments This work has been submitted to IFAC for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03722 2025-12-04 cs.NI 89%

Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks

无线网络中增强型强化学习教程:大型语言模型

Lingyi Cai, Wenjie Fu, Yuxi Huang, Ruichen Zhang, Yinqiu Liu, Jiawen Kang, Zehui Xiong, Tao Jiang, Dusit Niyato, Xianbin Wang, Shiwen Mao, Xuemin Shen

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文介绍了利用大型语言模型增强无线网络强化学习的方法,涵盖分类框架、角色功能及应用场景,探讨未来发展方向。

Comments 30 pages, 12 figures, survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03421 2025-12-04 cs.SE 89%

Exploring the Potential and Limitations of Large Language Models for Novice Program Fault Localization

探索大型语言模型在初级程序员故障定位中的潜力与局限性

Hexiang Xu, Hengyuan Liu, Yonghao Wu, Xiaolan Kang, Xiang Chen, Yong Liu

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文评估了LLM在初级程序员故障定位中的性能,发现高级模型在准确性上表现优异,但存在推理过度和计算成本高等问题。

Comments The paper has been accepted for publication in The Journal of Systems & Software

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09060 2025-12-04 cs.DB 89%

GenRewrite: Query Rewriting via Large Language Models

GenRewrite: 通过大语言模型实现查询重写

Jie Liu, Barzan Mozafari

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 GenRewrite 利用大语言模型实现查询重写,通过自然语言重写规则和反例引导技术提升查询优化效率。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03111 2025-12-04 cs.CR 85%

Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents

Les Dissonances:池中工具赋能LLM代理中的跨工具采集与污染

Zichuan Li, Jian Cui, Xiaojing Liao, Luyi Xing

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种新型威胁XTHP,揭示了LLM代理中多工具整合带来的安全风险,并通过Chord工具验证了75%的现实世界工具存在此漏洞。

Comments Network and Distributed System Security (NDSS) Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03272 2025-12-04 cs.AI 84%

When Do Symbolic Solvers Enhance Reasoning in Large Language Models?

符号求解器在大语言模型中增强推理何时有效?

Zhiyuan He, Dingmin Wang

机构 * University College London(伦敦大学学院) University of Oxford(牛津大学)

专题命中 推理与问题求解 :large language model(title);language model(title);分类 cs.AI

AI总结 本文研究符号求解器如何在大语言模型中增强推理,发现其在需要有限隐式推理但具有充足搜索空间的问题中表现优异,尤其在约束满足问题中提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03082 2025-12-04 cs.CL cs.AI 84%

Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation

通过推理依赖生成缓解LLM中的选择支持性偏见

Nan Zhuang, Wenshuo Wang, Lekai Qian, Yuxiao Wang, Boyu Cao, Qi Liu

机构 * School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过推理依赖生成缓解LLM中选择支持性偏见的方法,生成无偏推理数据并提升模型在决策任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23143 2025-12-04 cs.AI cs.LG cs.SY eess.SY 84%

MathBode: Measuring the Stability of LLM Reasoning using Frequency Response

MathBode: 通过频率响应测量大语言模型推理的稳定性

Charles L. Wang

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 MathBode通过频率响应分析大语言模型推理的稳定性,揭示系统性低通行为和相位滞后,提供可重复的动态诊断方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07506 2025-12-04 cs.DC cs.AI cs.CL cs.LG cs.SE 83%

Astra: A Multi-Agent System for GPU Kernel Performance Optimization

Astra:一种用于GPU内核性能优化的多智能体系统

Anjiang Wei, Tianran Sun, Yogesh Seenichamy, Hang Song, Anne Ouyang, Azalia Mirhoseini, Ke Wang, Alex Aiken

机构 * Stanford University(斯坦福大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University(南京大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 Astra是首个基于LLM的多智能体系统,用于优化GPU内核性能,通过协作生成高效内核并实现显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03360 2025-12-04 cs.CL 83%

From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation

从假设到前提:基于大语言模型的反向逻辑推理与选择性符号翻译

Qingchuan Li, Mingyue Cheng, Zirui Liu, Daoyu Wang, Yuting Zeng, Tongxuan Liu

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出HBLR框架,通过结合自信符号翻译与假设驱动反向推理,提升逻辑推理的准确性和效率。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18098 2025-12-04 cs.CL cs.AI 82%

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

无需搜索的规划:通过离线目标条件强化学习精炼前沿大语言模型

Joey Hong, Anca Dragan, Sergey Levine

机构 * UC Berkeley(伯克利大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 通过目标条件价值函数引导LLM推理,实现高效多轮交互规划,优于传统RL微调和提示方法。

Comments Published at NeurIPS 2025; 18 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03276 2025-12-04 cs.LG 81%

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

太迟回忆:多模态知识检索中双跳问题的解释

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, Neel Nanda

机构 * University of Oxford(牛津大学) McGill University(麦吉尔大学) Meta MATS

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文研究了多模态知识检索中双跳问题的影响,发现VLMs在处理视觉输入时过晚解决实体表示导致事实回忆性能下降,并提出通过修补和提示方法恢复性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05826 2025-12-04 cs.CV cs.AI cs.LG 81%

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

从像素到 prose:推进遥感多模态语言模型

Xintian Sun, Benji Peng, Charles Zhang, Fei Jin, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Xinyuan Song, Yichao Zhang

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Minnesota - Twin Cities(明尼苏达大学双城分校) Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 推理与问题求解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了多模态语言模型在遥感中的应用,分析了其技术基础、挑战及未来发展方向,强调了其在环境监测和灾害响应中的重要作用。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17701 2025-12-04 cs.CL cs.AI cs.LG 80%

Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs

探究偏差:一种多语言流水线用于生成、解决和评估LLM数学问题

Mariam Mahran, Katharina Simbeck

机构 * HTW Berlin University of Applied Sciences(柏林应用科学大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出多语言流水线生成、解决和评估数学问题,发现英语解决方案质量优于阿拉伯语,揭示语言偏差问题。

Comments Published in CEUR Workshop Proceedings, Vol. 4114, edu4AI'25: 2nd Workshop on Education for Artificial Intelligence, co-located with ECAI 2025, Bologna, Italy

Journal ref In: Proceedings of the 2nd International Workshop on Education for Artificial Intelligence (edu4AI 2025), CEUR Workshop Proceedings, Vol. 4114, Bologna, Italy, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03759 2025-12-04 cs.CL cs.AI cs.LG 75%

Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective

从序列层面视角出发的扩散大语言模型原理化强化学习

Jingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu, Jianwen Xie, Stefano Ermon, Yi Wu, Chongxuan Li

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Stanford University(斯坦福大学) Lambda, Inc(Lambda公司) Tsinghua University(清华大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出ESPO,一种基于序列级优化的原理化强化学习框架,用于提升扩散大语言模型的生成能力,通过ELBO作为似然代理,显著优于token级方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03180 2025-12-04 cs.MA cs.ET 75%

AGENTSAFE: A Unified Framework for Ethical Assurance and Governance in Agentic AI

AGENTSAFE:面向代理AI的统一伦理保障与治理框架

Rafflesia Khan, Declan Joyce, Mansura Habiba

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 AGENTSAFE提出一个统一的治理框架,通过风险分类转化为可操作的设计、运行时和审计控制,提供可衡量的预部署保障,并通过运行时治理和问责机制建立代理AI生态系统的信任。

Comments 12 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03874 2025-12-04 cs.RO cs.LG 74%

OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance

OmniDexVLG: 从视觉语言模型引导的抓取语义、分类法和功能可及性中学习灵巧抓取生成

Lei Zhang, Diwen Zheng, Kaixin Bai, Zhenshan Bing, Zoltan-Csaba Marton, Zhaopeng Chen, Alois Christian Knoll, Jianwei Zhang

机构 * University of Hamburg(汉堡大学) Agile Robots SE(敏捷机器人公司) Technical University of Munich(慕尼黑技术大学)

专题命中 推理与问题求解 :language model(title);分类 cs.LG

AI总结 OmniDexVLG通过多模态语义推理和统一视觉语言模型,实现从自然语言指令中生成多样且语义连贯的灵巧抓取。

Comments Project Website: https://sites.google.com/view/omnidexvlg, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00926 2025-12-04 cs.AI cs.CL 73%

LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory

大语言模型将自己定位为比人类更理性:通过博弈论测量AI自我意识的出现

Kyung-Hoon Kim

机构 * Gmarket Seoul, South Korea(韩国首尔Gmarket)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究通过博弈论框架测量大语言模型的自我意识,发现先进模型表现出比人类更高的理性认知。

Comments 19 pages, 6 figures, 28 models tested across 4,200 trials

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19306 2025-12-04 cs.AI cs.CL 73%

SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling

SETS:利用自我验证和自我校正以提高测试时间扩展

Jiefeng Chen, Jie Ren, Xinyun Chen, Chengrun Yang, Ruoxi Sun, Jinsung Yoon, Sercan Ö Arık

机构 * Google Cloud AI Research(谷歌云人工智能研究)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SETS通过结合并行与顺序技术,利用LLMs的自我改进能力,在无需模型训练的情况下提升测试时间扩展性能。

Comments Published in Transactions on Machine Learning Research (11/2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03233 2025-12-04 cs.CV 67%

Object Counting with GPT-4o and GPT-5: A Comparative Study

基于GPT-4o和GPT-5的对象计数:比较研究

Richard Füzesséry, Kaziwa Saleh, Sándor Szénási, Zoltán Vámossy

机构 * Software Engineering Institute, Obuda University, Budapest, Hungary(奥布达大学软件工程研究所) Doctoral School of Applied Informatics(应用信息科学博士学院) Applied Mathematics, Obuda University, Budapest, Hungary(应用数学,奥布达大学) John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary(约翰·冯·诺伊曼信息学院,奥布达大学) Faculty of Economics(经济学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

AI总结 本文比较了GPT-4o和GPT-5在零样本对象计数任务中的性能,展示了其在FSC-147和CARPK数据集上的表现。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03176 2025-12-04 cs.LG cs.AI 62%

Plantain: Plan-Answer Interleaved Reasoning

Plantain: 计划-答案交错推理

Anthony Liang, Jonathan Berant, Adam Fisch, Abhimanyu Goyal, Kalpesh Krishna, Jacob Eisenstein

机构 * Google DeepMind(谷歌DeepMind) University of Southern California(南加州大学)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI、cs.LG

AI总结 Plantain通过计划-思考-回答的交错推理方式,提升数学推理和编码任务的性能,同时减少响应时间。

详情

展开后加载摘要…

URL PDF HTML 收藏