arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-02 至 2026-03-02 共收录 180 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 41 篇

2508.18395 2026-03-02 cs.CL cs.AI 79%

Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning

潜在自一致性用于短答和长答推理中可靠多数集选择

Jungsuk Oh, Jay-Yoon Lee

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 LSC通过学习token嵌入选择语义一致的响应,有效提升短答和长答推理中的多数集选择可靠性,且计算开销低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03632 2026-03-02 cs.AI 77%

MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information

MITS: 通过点互信息增强大语言模型的树搜索推理

Jiaxi Li, Yucheng Shi, Xiao Huang, Jin Lu, Ninghao Liu

机构 * University of Georgia(佐治亚大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 MITS通过点互信息引导树搜索推理,提升大语言模型的推理性能和效率。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24532 2026-03-02 cs.CL 77%

DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance

DeepQuestion: 为评估LLM性能系统生成现实世界挑战

Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi, Heshaam Faili

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 DeepQuestion通过系统生成现实世界挑战,揭示LLM在复杂任务上的性能下降,强调需多样化评估以改进模型发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24241 2026-03-02 cs.IR cs.HC 75%

UXSim: Towards a Hybrid User Search Simulation

UXSim: 向混合用户搜索模拟迈进

Saber Zerhoudi, Michael Granitzer

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 UXSim 通过整合传统模拟器与适应性 LLM 代理,提供更准确的用户行为模拟和可解释的验证方法。

Journal ref Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25), November 10--14, 2025, Seoul, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23635 2026-03-02 cs.HC 75%

When LLMs Help -- and Hurt -- Teaching Assistants in Proof-Based Courses

当LLMs助益与损害助教在证明基础课程中的表现

Romina Mahinpei, Sofiia Druchyna, Manoel Horta Ribeiro

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文研究了LLM在证明基础课程中助教评分与反馈中的作用,发现LLM与助教在评分上存在分歧,但其反馈对存在重大错误的提交仍具帮助。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23623 2026-03-02 cs.NI 75%

Toward E2E Intelligence in 6G Networks: An AI Agent-Based RAN-CN Converged Intelligence Framework

迈向6G网络的端到端智能:一种基于AI代理的RAN-CN融合智能框架

Youbin Han, Haneul Ko, Namseok Ko, Tarik Taleb, Yan Chen

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出一种基于AI代理的RAN-CN融合智能框架,利用大语言模型和ReAct范式,实现跨域网络的统一智能控制,提升网络适应性和泛化能力。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21135 2026-03-02 cs.RO cs.AI cs.CV 74%

SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation

SocialNav: 为社交感知具身导航训练的人类启发式基础模型

Ziyi Chen, Yingnan Guo, Zedong Chu, Minghua Luo, Yanfen Shen, Mingchao Sun, Junjun Hu, Shichao Xie, Kuan Yang, Pei Shi, Zhining Gu, Lu Liu, Honglin Han, Xiaolong Wu, Mu Xu, Yu Zhang, Ning Guo

机构 * Amap, Alibaba Group, China(阿里集团阿地图,中国) Zhejiang University, China(浙江大学,中国)

专题命中 推理与问题求解 :foundation model(title);分类 cs.AI

AI总结 SocialNav通过分层架构和多阶段训练,实现了社交感知导航,显著提升了导航性能和社会合规率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23945 2026-03-02 cs.CV cs.AI cs.MM 70%

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

PointCoT: 一种用于显式3D几何推理的多模态基准

Dongxu Zhang, Yiding Sun, Pengcheng Li, Yumou Liu, Hongqiang Lin, Haoran Xu, Xiaoxuan Mu, Liang Lin, Wenbiao Yan, Ning Yang, Chaowei Fang, Juanjuan Zhao, Jihua Zhu, Conghui He, Cheng Tan

机构 * Xi'an Jiaotong University(西安交通大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Institute of Automation, CASIA(中国科学院自动化研究所) Taiyuan University of Technology(太原理工大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 PointCoT通过显式链式推理提升3D几何推理能力,提出多模态基准和双流架构,实现对3D点云的高精度理解与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22543 2026-03-02 cs.LG 70%

FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning

FAPO:面向高效可靠推理的缺陷感知策略优化

Yuyang Ding, Chi Zhang, Juntao Li, Haibin Lin, Min Zhang

机构 * Soochow University(苏州大学) ByteDance Seed(字节跳动种子)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 FAPO通过引入缺陷感知机制,优化策略以提升大语言模型在高效可靠推理中的性能,提高结果正确性与训练稳定性。

Comments ICLR 2026. Project page: https://fapo-rl.github.io/; Infra Doc: https://verl.readthedocs.io/en/latest/advance/reward_loop.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21048 2026-03-02 cs.CV cs.AI 70%

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

Veritas:通过模式感知推理实现通用的深度伪造检测

Hao Tan, Jun Lan, Zichang Tan, Ajian Liu, Chuanbiao Song, Senyuan Shi, Huijia Zhu, Weiqiang Wang, Jun Wan, Zhen Lei

机构 * School of Advanced Interdisciplinary Sciences (SAIS), University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Ant Group(蚂蚁集团) Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Veritas通过模式感知推理,基于多模态大语言模型实现通用深度伪造检测,提升对未知伪造技术和数据领域的检测能力。

Comments ICLR 2026 Oral. Project: https://github.com/EricTan7/Veritas

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23959 2026-03-02 cs.CV 67%

Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought

通过图像作为连续动作进行思考:数值视觉链式推理

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Chalmers University of Technology(楚克理工大学) University of Gothenburg(哥德堡大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

AI总结 NV-CoT通过将图像推理动作空间扩展为连续欧几里得空间,提升MLLMs的定位精度和回答准确性,同时加速训练收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23952 2026-03-02 cs.CV 67%

CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering

CC-VQA: 一种考虑冲突和相关性的方法,用于缓解基于知识的视觉问答中的知识冲突

Yuyang Hong, Jiaqi Gu, Yujin Lou, Lubin Fan, Qi Yang, Ying Wang, Kun Ding, Yue Wu, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, UCAS(中国科学技术大学人工智能学院) MAIS, Institute of Automation(自动化研究所MAIS) Alibaba Cloud Computing(阿里云计算)

专题命中 推理与问题求解 :language model(abstract);prompting(abstract)

AI总结 CC-VQA提出了一种无训练、考虑冲突和相关性的方法,通过视觉中心的上下文冲突推理和相关性引导的编码解码,提升基于知识的视觉问答的准确率。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20812 2026-03-02 cs.CV cs.AI cs.CL 62%

Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation

小草稿,大结论:通过推测进行信息密集的视觉推理

Yuhan Liu, Lianhui Qin, Shengjie Wang

机构 * New York University(纽约大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 推理与问题求解 :language model(abstract);分类 cs.CL、cs.AI

AI总结 SV通过结合多个轻量级草稿专家和大型结论模型,提升信息密集视觉推理的效率和准确性。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23546 2026-03-02 cs.CL cs.AI 62%

Humans and LLMs Diverge on Probabilistic Inferences

人类与大语言模型在概率推断上存在分歧

Gaurav Kamath, Sreenath Madathil, Sebastian Schuster, Marie-Catherine de Marneffe, Siva Reddy

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所) University of Vienna(维也纳大学) FNRS – UCLouvain(佛罗伦萨-鲁汶大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 推理与问题求解 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文通过ProbCOPA数据集研究人类与大语言模型在概率推断上的差异,发现模型在非确定性推理上表现不足,揭示了人类与LLM在推理模式上的根本差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22453 2026-03-02 cs.CL 57%

Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads

通过检索-过渡头桥接潜在推理与目标语言生成

Shaswat Patel, Vishvesh Trivedi, Yue Han, Yihuai Hong, Eunsol Choi

专题命中 推理与问题求解 :language model(abstract);分类 cs.CL

AI总结 本研究通过引入检索-过渡头(RTH)揭示多语言模型中负责目标语言生成的关键注意力机制,并通过实验验证其在链式推理中的重要性。

Comments In the paper, there are still many statements that are unclear and lack sufficient justification. Since it is difficult for us to estimate how much time would be required to properly revise and correct these issues, we would like to request a withdrawal of the paper in this moment. Thank you!

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06940 2026-03-02 cs.DB cs.AI 57%

VISTA: Knowledge-Driven Vessel Trajectory Imputation with Repair Provenance

VISTA: 基于知识的船舶轨迹补全与修复溯源

Hengyu Liu, Tianyi Li, Haoyu Wang, Kristian Torp, Tiancheng Zhang, Yushuai Li, Christian S. Jensen

机构 * Department of Computer Science, Aalborg University, Denmark(计算机科学系,奥胡斯大学,丹麦) School of Computer Science and Engineering, Northeastern University, Shenyang, China(计算机科学与工程学院,东北大学,沈阳,中国)

专题命中 推理与问题求解 :LLM(abstract);分类 cs.AI

AI总结 VISTA通过基于知识的可解释方法实现船舶轨迹补全,生成修复溯源以提升下游决策的可信度和效率。

Comments 24 pages, 14 figures, 4 algorithms, 8 tables. Code available at https://github.com/hyLiu1994/VISTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02814 2026-03-02 cs.AI 57%

Radiologist Copilot: An Agentic Framework Orchestrating Specialized Tools for Reliable Radiology Reporting

放射科助手:一种协调专用工具的代理框架以实现可靠的放射学报告

Yongrui Yu, Zhongzhen Huang, Linjie Mu, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学) SenseTime Research, Shanghai, China(商汤研究) Shanghai Innovation Institute, Shanghai, China(上海创新研究院)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

AI总结 放射科助手通过协调专用工具实现全面的放射学报告流程,提升报告准确性和临床标准符合度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04506 2026-03-02 cs.CL 57%

Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathways

在放射学报告中建模临床不确定性:从显式不确定性标记到隐式推理路径

Paloma Rabaey, Jong Hak Moon, Jung-Oh Lee, Min Gwan Kim, Hangyul Yoon, Thomas Demeester, Edward Choi

专题命中 推理与问题求解 :LLM(abstract);分类 cs.CL

AI总结 本研究提出Lunguage++基准,通过显式和隐式不确定性建模方法,提升放射学报告中不确定性分析的准确性和应用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25249 2026-03-02 cs.RO cs.AI 57%

BEV-VLM: Trajectory Planning via Unified BEV Abstraction

BEV-VLM:通过统一的鸟瞰抽象进行轨迹规划

Guancheng Chen, Sheng Yang, Tong Zhan, Jian Wang

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

AI总结 BEV-VLM通过统一的BEV抽象与VLMs实现高精度轨迹规划,实验显示其在准确性上比现有方法提升53.1%并实现无碰撞。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 评测与基准 30 篇

2505.19764 2026-03-02 cs.LG cs.AI 89%

Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows

多视角编码器在基于大语言模型的代理工作流性能预测中的应用

Patara Trirat, Wonyong Jeong, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本文提出Agentic Predictor,通过多视角编码技术与跨领域预训练,实现高效准确的代理工作流性能预测,提升LLM代理系统设计效率。

Comments ICLR 2026, Project Page: https://deepauto-ai.github.io/agentic-predictor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03716 2026-03-02 cs.CL cs.LG hep-th 88%

FeynTune: Large Language Models for High-Energy Theory

FeynTune:用于高能理论的大型语言模型

Paul Richmond, Prarit Agarwal, Borun Chowdhury, Vasilis Niarchos, Constantinos Papageorgakis

机构 * Centre for Theoretical Physics, Department of Physics and Astronomy, Queen Mary University of London(伦敦大学玛丽女王学院理论物理中心,物理与天文学系) ITCP & CCTP, Department of Physics, University of Crete(希腊克里特大学物理系ITCP与CCTP) Meta Amazon(亚马逊公司)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 FeynTune通过微调Llama-3.1模型,为高能理论物理开发了专门的大型语言模型,并在多个领域数据集上进行了性能比较。

Comments 16 pages; v2: Human evaluation discussion updated, additional training hyperparameters and inference settings included and references added

Journal ref Mach. Learn.: Sci. Technol. 7 025012 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23980 2026-03-02 cs.CV 88%

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

Venus: 多模态大语言模型在审美指导与裁剪中的基准测试与赋能

Tianxiang Du, Hulingxiao He, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究所)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

AI总结 Venus通过两阶段框架提升多模态大语言模型的审美指导与裁剪能力,实现可解释的审美优化。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19966 2026-03-02 cs.CL cs.AI 87%

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Dhati+: 针对阿拉伯语主观性评估的微调大型语言模型

Slimane Bellaouar, Attia Nehar, Soumia Souffi, Mounia Bouameur

机构 * Dept. of Mathematics and Computer Science(数学与计算机科学系) Lab. des Mathématiques et Sciences Appliquées (LMSA)(应用数学与科学实验室(LMSA)) Exact Sciences and Computer Science Faculty(精确科学与计算机科学系) Lab. d’Informatique et Mathématiques (LIM)(计算机与数学实验室(LIM))

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL、cs.AI

AI总结 本文提出了一种针对阿拉伯语主观性评估的微调大型语言模型方法,通过构建AraDhati+数据集并结合集成决策方法,实现了97.79%的高准确率。

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23373 2026-03-02 cs.AI cs.CL cs.IR 86%

An Agentic LLM Framework for Adverse Media Screening in AML Compliance

面向反面媒体筛查的代理LLM框架用于反洗钱合规

Pavel Chernakov, Sasan Jafarnejad, Raphaël Frank

机构 * Interdisciplinary Centre for Security, Reliability and Trust(安全、可靠与信任跨学科中心)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于LLM与RAG的代理框架,用于自动化反面媒体筛查,通过多步骤流程实现高风险个体的准确识别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25662 2026-03-02 cs.HC cs.AI 85%

User Misconceptions of LLM-Based Conversational Programming Assistants

基于大语言模型的对话式编程助手中的用户误解

Gabrielle O'Brien, Antonio Pedro Santos Alves, Sebastian Baltes, Grischa Liebel, Mircea Lungu, Marcos Kalinowski

机构 * University of Michigan(密歇根大学) Pontifical Catholic University of Rio de Janeiro(里约热内卢天主教大学) Heidelberg University(海德堡大学) Reykjavik University(雷克雅未克大学) IT University of Copenhagen(哥本哈根IT大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究揭示了用户在使用基于大语言模型的对话式编程助手时可能存在的误解,强调了工具在明确自身能力及通过实证研究澄清编程需求的重要性。

Comments Accepted to the Journal Ahead Workshop at ICSE '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22291 2026-03-02 cs.LG cs.AI cs.CR 84%

Manifold of Failure: Behavioral Attraction Basins in Language Models

失败的流形:语言模型中的行为吸引盆地

Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala, Idan Habler, Ammar Al-Kahfah, Ken Huang, Blake Gatto

机构 * Amazon Web Services(亚马逊网络服务) Cisco(思科) Shrewd Security(精明安全)

专题命中 评测与基准 :language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于MAP-Elites的方法,用于系统地映射语言模型的失败流形,揭示不同模型的拓扑结构和漏洞分布。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23676 2026-03-02 cs.CV 80%

Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering

通过语义解耦潜在引导抑制放射科报告生成中的先验比较幻觉

Ao Li, Rui Liu, Mingjie Li, Sheng Liu, Lei Wang, Xiaodan Liang, Lina Yao, Xiaojun Chang, Lei Xing

机构 * University of New South Wales(新南威尔士大学) Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,技术悉尼大学) Stanford University(斯坦福大学) School of Computing and Information Technology of University of Wollongong Australia(沃林根澳大利亚大学计算与信息科技学院) Sun Yat-sen University(中山大学) University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出语义解耦潜在引导方法,通过正交化技术减少放射科报告生成中的历史幻觉,提升临床准确性与报告忠实度。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04026 2026-03-02 cs.HC 80%

VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking

VeriWeb: 可验证的长链网络基准测试用于代理信息检索

Shunyu Liu, Minghao Liu, Huichi Zhou, Zhenyu Cui, Yang Zhou, Yuhao Zhou, Jialiang Gao, Heng Zhou, Yunhao Yang, Wendong Fan, puzhen zhang, Ge Zhang, Jiajun Shi, Weihao Xuan, Jiaxing Huang, Shuang Luo, Fang Wu, Heli Qi, Qingcheng Zeng, Junjie Wang, Aosong Feng, Jindi Lv, Sicong Jiang, Ziqi Ren, Wangchunshu Zhou, Zhenfei Yin, Wenlong Zhang, Guohao Li, Wenhao Yu, Lei Ma, Lei Bai, Qunshu Lin, Mingli Song, Dacheng Tao

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 VeriWeb是一个用于评估和开发网络代理的可验证长链网络基准测试,通过长链复杂性和子任务级可验证性设计,揭示了处理长时间网络任务的性能差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23678 2026-03-02 cs.CV cs.LG 79%

Any Model, Any Place, Any Time: Get Remote Sensing Foundation Model Embeddings On Demand

任何模型、任何地点、任何时间:获取遥感基础模型嵌入的按需服务

Dingqi Ye, Daniel Kiv, Wei Hu, Jimeng Shi, Shaowen Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 rs-embed库通过统一接口实现任意模型、地点和时间的遥感嵌入按需获取,提升模型使用效率与公平比较能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20640 2026-03-02 cs.AI cs.LG 79%

CoMind: Towards Community-Driven Agents for Machine Learning Engineering

CoMind:面向机器学习工程的社区驱动代理

Sijie Li, Weiwei Sun, Shanda Li, Ameet Talwalkar, Yiming Yang

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 CoMind通过多代理系统和迭代并行探索机制,在Kaggle竞赛中实现了优于人类竞争者的性能,展示了社区驱动代理在机器学习工程中的潜力。

Comments ICLR 2026. Code available at https://github.com/comind-ml/CoMind

详情

展开后加载摘要…

URL PDF HTML 收藏