arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Stanford University(斯坦福大学)

2026-04-30 至 2026-04-30 共收录 8
2604.22750 2026-04-30 cs.CL cs.AI cs.CY cs.HC cs.SE

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

AI代理如何花费你的钱?分析和预测代理编码任务中的令牌消耗

Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei

机构 * University of Michigan(密歇根大学) Stanford University(斯坦福大学) All Hands AI Google Deepmind(谷歌DeepMind) Microsoft AI(微软AI) Massachusetts Institute of Technology(麻省理工学院)

AI总结 本文研究了代理编码任务中令牌消耗模式,发现代理任务消耗远高于代码推理和代码聊天,且模型在任务执行前难以准确预测自身令牌使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26479 2026-04-30 stat.ME cs.LG

Recipes for Calibration Checks in Safety-Critical Applications

安全关键应用中的校准检查配方

Romeo Valentin

机构 * Stanford University(斯坦福大学)

AI总结 本文提出了一种校准检查框架,用于验证预测系统分布特性,通过模块化流程支持多种应用场景,展示了在天气预测和机器人姿态估计中的应用。

Comments 36 pages, 22 figures. Manuscript prepared with Typst

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26211 2026-04-30 cs.AI cs.LG

OMEGA: Optimizing Machine Learning by Evaluating Generated Algorithms

OMEGA:通过评估生成算法优化机器学习

Jeremy Nixon, Annika Singh

机构 * Infinity Artificial Intelligence Institute(无限人工智能研究所) Stanford University(斯坦福大学)

AI总结 OMEGA框架结合结构化元提示工程与可执行代码生成,创建新型ML分类器,在20个基准数据集上超越scikit-learn基线。

Comments ICLR 2026: Workshop on AI with Recursive Self-Improvement

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25982 2026-04-30 cs.LG cs.AI cs.CY cs.ET

Open Problems in Frontier AI Risk Management

前沿人工智能风险管理中的开放问题

Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, Robert Trager

机构 * Oxford Martin AI Governance Initiative, University of Oxford(牛津大学人工智能治理倡议) MIT Computer Science and Artificial Intelligence Laboratory, MIT(麻省理工学院计算机科学与人工智能实验室) MIT Future Tech(麻省理工学院未来技术) Stanford University(斯坦福大学) Governance and Responsible AI Lab, Purdue University(普渡大学治理与负责任的人工智能实验室) University of Toronto(多伦多大学) Mercatus Center, George Mason University(乔治·马歇尔大学麦卡锡中心) Vilnius University(维尔纽斯大学) Vijil SaferAI AI Standards Lab(人工智能标准实验室) The Future Society(未来社会) Concordia AI(康科德人工智能) Pivotal Research Center for AI Risk Management & Alignment(人工智能风险管理和对齐中心) UC Berkeley Center for Long-Term Cybersecurity(伯克利大学长期网络安全中心) Independent(独立)

AI总结 本文探讨前沿人工智能风险管理中的核心问题,通过文献综述识别未解决的挑战,并分类问题类型以指导未来研究与治理。

Comments 81 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23575 2026-04-30 cs.CY cs.CL cs.LG

The Collapse of Heterogeneity in Silicon Philosophers

硅哲学家中的异质性崩溃

Yuanming Shi, Andreas Haupt

机构 * Adobe Inc.(Adobe公司) Stanford University(斯坦福大学)

AI总结 研究发现硅样本在哲学领域系统性地压缩了异质性,语言模型过度相关哲学判断,产生人工共识,影响对齐、评估和硅样本替代人类判断的使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04325 2026-04-30 cs.CL cs.AI cs.CV cs.LG cs.MM

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

超越排行榜:重新思考大型语言模型的医疗基准

Wenting Chen, Guo Yu, Yiu-Fai Cheung, Meidan Ding, Jie Liu, Zizhan Ma, Wenxuan Wang, Linlin Shen

机构 * Stanford University(斯坦福大学) Shenzhen University(深圳大学) The Chinese University of Hong Kong(香港中文大学) City University of Hong Kong(香港城市大学) Renmin University of China(中国人民大学)

AI总结 本文提出MedCheck框架,用于评估医疗领域大型语言模型的基准,揭示现有基准在临床相关性、数据完整性及安全性方面的系统性问题。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01544 2026-04-30 cs.LG

MARVIS: Modality Adaptive Reasoning over VISualizations

MARVIS:基于可视化模态自适应推理

Benjamin Feuer, Lennart Purucker, Oussama Elachqar, Chinmay Hegde

机构 * Stanford University(斯坦福大学) Prior Labs(Prior实验室) New York University(纽约大学)

AI总结 MARVIS通过将潜在嵌入空间转换为可视化表示,并利用VLM的空间和细粒度推理能力,实现跨视觉、音频、生物和表格域的预测,以3B参数模型在多个领域超越Gemini 2.0。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20876 2026-04-30 cs.CL

Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine

少做决定,多进行沟通:关于医学领域端到端事实核查构建效度的探讨

Sebastian Joseph, Lily Chen, Barry Wei, Michael Mackert, Iain J. Marshall, Paul Pu Liang, Ramez Kouzy, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Stanford University(斯坦福大学) Indiana University School of Medicine(印第安纳大学医学院) King’s College London(伦敦国王学院) Massachusetts Institute of Technology(麻省理工学院) The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Northeastern University(东北大学)

AI总结 本文探讨医学领域端到端事实核查系统的构建效度,指出其在连接现实声明与科学证据、处理模糊声明及主观真实性标签方面的挑战,主张将其视为互动沟通问题。

Comments ACL 2026 Findings camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏