arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

2026-01-30 至 2026-01-30 共收录 9
2510.25889 2026-01-30 cs.LG

$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

$π_\ exttt{RL}$: 流基于视觉-语言-动作模型的在线强化学习微调

Kang Chen, Zhihao Liu, Tonghe Zhang, Zhen Guo, Si Xu, Hao Lin, Hongzhi Zang, Xiang Li, Quanlu Zhang, Zhaofei Yu, Guoliang Fan, Tiejun Huang, Yu Wang, Chao Yu

机构 * Tsinghua University(清华大学) Peking University(北京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Carnegie Mellon University(卡内基梅隆大学) Infinigence AI Zhongguancun Academy(中关村学院)

AI总结 本文提出$π_\ exttt{RL}$方法,通过流噪声和流SDE技术,解决大规模流基于VLA模型中强化学习微调的挑战,提升模型在分布内和分布外任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07042 2026-01-30 cs.HC cs.AI cs.CL cs.CY

Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions

Minion:一种技术探测器,用于探索用户如何与AI伙伴协商有害的价值冲突

Xianzhe Fan, Qing Xiao, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen

机构 * The University of Hong Kong(香港大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Tsinghua University(清华大学) Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

AI总结 Minion通过探测用户与AI伙伴协商有害价值冲突的过程,揭示了设计中需平衡用户责任与AI安全的挑战。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21147 2026-01-30 cs.LG

Smooth Dynamic Cutoffs for Machine Learning Interatomic Potentials

为机器学习互原子势设计平滑的动态截断

Kevin Han, Haolin Cong, Bowen Deng, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学) Massachuessets Institute of Technology(麻省理工学院)

AI总结 本研究提出动态截断方法,通过诱导原子图稀疏性,显著降低MLIPs的内存消耗和推理时间,同时保持模拟稳定性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21082 2026-01-30 cs.LG cs.AI

LOCUS: Low-Dimensional Model Embeddings for Efficient Model Exploration, Comparison, and Selection

LOCUS:低维模型嵌入用于高效模型探索、比较和选择

Shivam Patel, William Cocke, Gauri Joshi

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 LOCUS通过低维嵌入高效实现模型探索、比较和选择,减少查询样本需求并提升模型相似性判断能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21051 2026-01-30 cs.AI cs.CR cs.LG

Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report

Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B 技术报告

Zhuoran Yang, Ed Li, Jianliang He, Aman Priyanshu, Baturay Saglam, Paul Kassianik, Sajana Weerawardhena, Anu Vellore, Blaine Nelson, Neusha Javidnia, Arthur Goldblatt, Fraser Burch, Avi Zohary, Assaf Eisenman, Mahdi Sabbaghi, Supriti Vijay, Rahim Dharssi, Dhruv Kedia, Kojin Oshiba, Yaron Singer, Amin Karbasi

机构 * Foundation AI–Cisco Systems Inc.(Foundation AI–Cisco系统公司) Yale University(耶鲁大学) University of California, San Diego(加州大学圣地亚哥分校) University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B 是首个开源的网络安全专用推理模型,通过结合监督微调和强化学习训练,实现了在网络安全任务上的高性能表现。

Comments 31 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21049 2026-01-30 cs.AI

QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation

QUARK:通过查询锚定聚合实现非忠实查询下的稳健检索

Rita Qiuran Lyu, Michelle Manqiao Wang, Lei Shi

机构 * University of California, Berkeley(加州大学伯克利分校) Carnegie Mellon University(卡内基梅隆大学) Adobe Research(Adobe研究)

AI总结 QUARK通过建模查询不确定性与锚定聚合,提升非忠实查询下的检索鲁棒性与性能。

Comments 11 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22467 2026-01-30 cs.LG

GLUE: Gradient-free Learning to Unify Experts

GLUE: 无梯度学习以统一专家

Jong-Ik Park, Shreyas Chaudhari, Srinivasa Pranav, Carlee Joe-Wong, José M. F. Moura

机构 * Electrical and Computer Engineering, Carnegie Mellon University(卡内基梅隆大学电气与计算机工程系) Computer Engineering, Carnegie Mellon University(卡内基梅隆大学计算机工程系)

AI总结 GLUE通过无梯度两点SPSA方法统一专家模型,提升目标领域测试准确率达8.5%-9.1%

Comments To be published at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21691 2026-01-30 cs.LG math.ST stat.TH

On Uncertainty Calibration for Equivariant Functions

关于等变函数的不确定性校准

Edward Berman, Jacob Ginesin, Marco Pacini, Robin Walters

机构 * Department of Mathematics, Northeastern University(东北大学数学系) Carnegie Mellon University(卡内基梅隆大学) University of Trento & Fondazione Bruno Kessler(特伦托大学及布鲁诺·凯斯勒基金会) Khoury College of Computer Sciences, Northeastern University(东北大学计算机科学学院) Geometric Learning Lab(几何学习实验室)

AI总结 本文研究了等变函数与不确定性校准之间的关系,通过理论分析和实验验证,揭示了对称性不匹配对模型校准的影响。

Comments Published in Transactions on Machine Learning Research (TMLR). Code is available at https://github.com/EdwardBerman/EquiUQ . Excited to share this paper, comments welcome :D

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15711 2026-01-30 cs.LG stat.ME stat.ML

Towards Identifiable Latent Additive Noise Models

朝着可识别的潜在加性噪声模型迈进

Yuhang Liu, Zhen Zhang, Dong Gong, Erdun Gao, Biwei Huang, Mingming Gong, Anton van den Hengel, Kun Zhang, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) School of Computer Science and Engineering, The University of New South Wales(计算机科学与工程学院,新南威尔士大学) Halicioğlu Data Science Institute, University of California San Diego(Halicioğlu数据科学研究所,加州大学圣地亚哥分校) School of Mathematics and Statistics, The University of Melbourne(数学与统计学学院,墨尔本大学) Department of Philosophy, Carnegie Mellon University(哲学系,卡内基梅隆大学)

AI总结 本文提出了一种更通用的框架,通过约束函数类来建立部分识别性结果,并开发了灵活的学习方法,用于学习潜在因果表示。

详情

展开后加载摘要…

URL PDF HTML 收藏