arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Transactions on Machine Learning Research · 期刊 · Machine Learning

共收录 1859
2605.25966 2026-05-26 cs.LG cs.CL stat.ML

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training

在小于100M参数量化感知训练中映射调度策略与位宽边界

Christian Brandt Thomassen

机构 * Dwarf A/S(Dwarf公司)

AI总结 通过大规模实验研究子100M参数解码器语言模型中,量化感知训练的最佳学习率调度是否依赖于位宽,发现INT6 QAT无需不同调度,INT4在50M以上需wd33调度,以下则噪声主导。

Comments 20 pages, 6 figures, 4 tables. 1345 training runs total (720 + 625). Submitted for review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22322 2026-05-26 cs.LG

A Closer Look on Memorization in Tabular Diffusion Model: A Data-Centric Perspective

表格扩散模型中记忆化的深入探究:以数据为中心的观点

Zhengyu Fang, Zhimeng Jiang, Huiyuan Chen, Xiaoge Zhang, Kaiyu Tang, Xiao Li, Jing Li

机构 * Department of Computer and Data Sciences(计算机与数据科学系) Case Western Reserve University(凯斯西储大学) Department of Computer Science & Engineering(计算机科学与工程系) Texas A&M University(德克萨斯大学) Department of Biochemistry(生物化学系) Center for RNA Science and Therapeutics(RNA科学与治疗中心) Department of Biomedical Engineering(生物医学工程系)

AI总结 本文首次从数据角度研究表格扩散模型中的记忆化动态,通过量化每个真实样本的记忆化程度,发现少数样本贡献了大部分泄露,并提出两阶段缓解方法DynamicCut。

Comments Published in Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01397 2026-05-26 cs.LG cs.AI cs.NA math.NA

Message-Passing GNNs Fail to Approximate Sparse Triangular Factorizations

消息传递GNN无法近似稀疏三角分解

Vladislav Trifonov, Ekaterina Muravleva, Ivan Oseledets

机构 * AIC, Skoltech(斯克里普金技术大学人工智能中心) Skoltech AI4S Center(斯克里普金技术大学AI4S中心) Sberbank of Russia(俄罗斯储蓄银行) AIRI

AI总结 本文通过理论和实验证明,消息传递图神经网络在逼近稀疏三角分解时存在根本性局限,需要超越消息传递的架构创新。

Comments Camera-ready version published in Transactions on Machine Learning Research

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00511 2026-05-26 cs.LG math.OC

Partition of Unity Neural Networks for Interpretable Classification with Explicit Class Regions

用于可解释分类的单元划分神经网络及显式类别区域

Akram Aldroubi

机构 * Department of Mathematics(数学系)

AI总结 提出单元划分神经网络(PUNN),通过直接学习满足和为1的非负函数来定义类别概率,无需softmax层,实现可解释分类并证明其稠密性,实验表明在保持精度的同时大幅减少参数。

Comments v2: substantially revised; under review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20738 2026-05-26 cs.LG cs.DC eess.SP math.OC stat.ML

SA-PEF: Step-Ahead Partial Error Feedback for Efficient Federated Learning

SA-PEF:用于高效联邦学习的前瞻部分误差反馈

Dawit Kiros Redie, Reza Arablouei, Stefan Werner

机构 * Department of Electronic Systems, Norwegian University of Science and Technology (NTNU)(挪威科学技术大学电子系统系) Department of Information and Communications Engineering, Aalto University(阿尔托大学信息与通信工程系) CSIRO’s Data61(CSIRO数据61)

AI总结 提出SA-PEF方法,通过结合前瞻校正和部分误差反馈,在非IID数据和部分客户端参与下加速联邦学习收敛,并理论证明其收敛速率与Fed-SGD相当。

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16763 2026-05-26 cs.CV

Flow Matching for Probabilistic Monocular 3D Human Pose Estimation

基于流匹配的概率单目3D人体姿态估计

Cuong Le, Pavlo Melnyk, Bastian Wandt, Mårten Wadenbäck

机构 * Department of Electrical Engineering(电气工程系) Linköping University(林雪平大学) Independent researcher(独立研究者)

AI总结 提出FMPose方法,利用流匹配生成模型从2D关键点学习3D姿态分布,通过图卷积网络建模2D提升条件,在保持精度的同时显著提升推理速度。

Comments 12 pages, 2 figures, 8 tables, accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11428 2026-05-26 cs.LG

Diagnosing Failure Modes of Neural Operators Across Diverse PDE Families

诊断不同PDE族中神经算子的失败模式

Lennon Shikhman

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文提出一个标准化压力测试框架,通过在不同PDE族上测试FNO、DeepONet和CNO三种架构,发现分布内准确率不能可靠预测鲁棒性,且失败模式依赖于架构和PDE族的组合。

Comments Published in Transactions on Machine Learning Research. 17 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11379 2026-05-26 stat.ML cs.LG math.ST stat.TH

Some Robustness Properties of Label Cleaning

标签清理的一些鲁棒性性质

Chen Cheng, John Duchi

AI总结 本文证明,依赖聚合标签(例如从噪声响应中提炼的标签信息)的学习过程具有数据清理无法实现的鲁棒性,体现在风险一致性、模型误设下的收敛性等方面。

Comments 41 pages, 3 figures. Accepted to Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03777 2026-05-26 cs.CV cs.LG

A Greedy Hierarchical Approach to Whole-Network Filter-Pruning in CNNs

一种面向CNN全网络滤波器剪枝的贪婪层次方法

Kiran Purohit, Anurag Reddy Parvathgari, Sourangshu Bhattacharya

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology, Kharagpur, India(印度理工学院,Khargpur,印度)

AI总结 提出一种基于线性近似的两层层次化贪婪剪枝算法,通过低层滤波器选择和全局剪枝准则高效剪枝,在多个网络上优于现有方法。

Comments Accepted in TMLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02416 2026-05-26 cs.LG stat.ML

Relative Translation Invariant Wasserstein Distance

相对平移不变Wasserstein距离

Binshuai Wang, Qiwei Di, Ming Yin, Mengdi Wang, Quanquan Gu, Peng Wei

机构 * Department of Computer Science(计算机科学系) George Washington University(乔治华盛顿大学) University of California, Los Angeles(加州大学洛杉矶分校) Department of Electrical and Computer Engineering(电气与计算机工程系) Princeton University(普林斯顿大学)

AI总结 受Bures距离启发,提出相对平移不变Wasserstein距离RW_p,证明其度量性质,并设计双层算法计算离散分布间的RW_p距离,当p=2时提出RW_2-LP和RW_2-Sinkhorn算法以提高数值稳定性,实验验证了算法在减少数值误差和实际雷暴模式检索中的有效性。

Comments Accepted by Transactions on Machine Learning Research (TMLR). Final accepted version. The implementation is publicly available at \url{https://github.com/DRKWang/rw_metric}

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13663 2026-05-25 cs.AI cs.LG

Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints

具有软实体约束的知识图谱交互式查询回答

Daniel Daza, Alberto Bernardi, Luca Costabello, Christophe Gueret, Masoud Mansoury, Michael Cochez, Martijn Schut

机构 * Translational AI Laboratory, Department of Laboratory Medicine(转化人工智能实验室,实验室医学系) Amsterdam University Medical Center, Vrije Universiteit Amsterdam(阿姆斯特丹大学医学中心,伏里埃大学阿姆斯特丹) Accenture Labs(埃森哲实验室) Delft University of Technology(代尔夫特理工大学) ELLIS Institute Finland & Abo Akademi University, Turku, Finland & Elsevier Discovery Lab, Amsterdam(芬兰ELLIS研究所 & 阿博阿卡迪米大学,图尔库,芬兰 & 埃西弗尔发现实验室,阿姆斯特丹)

AI总结 针对知识图谱查询中存在的模糊或上下文依赖的软约束问题,提出两种轻量级方法,通过调整查询答案分数来融入软约束,保持原始排序结构,并在扩展基准上验证了性能。

Comments Accepted in Transactions on Machine Learning Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18000 2026-05-25 cs.LG cs.AI q-bio.PE

Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

空间流行病模拟中的奖励工程:个体行为学习的强化学习平台

Radman Rakhshandehroo, Daniel Coombs

机构 * Department of Computer Science University of British Columbia(计算机科学系,不列颠哥伦比亚大学) Department of Mathematics and Institute of Applied Mathematics University of British Columbia(数学系和应用数学研究所,不列颠哥伦比亚大学)

AI总结 提出ContagionRL平台,通过奖励函数设计系统评估空间流行病模拟中个体行为学习策略,发现势场奖励方法能有效提升非药物干预依从性和空间规避策略。

Comments 38 pages, 15 figures and 18 tables; Accepted to TMLR. OpenReview: https://openreview.net/forum?id=yPEASsx3hk

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03508 2026-05-25 cs.LG

D2 Actor Critic: Diffusion Actor Meets Distributional Critic

D2 Actor Critic: 扩散演员遇上分布式评论家

Lunjun Zhang, Shuo Han, Hanrui Lyu, Bradly C Stadie

机构 * Department of Computer Science, University of Toronto(计算机科学系,多伦多大学) Department of Statistics, Northwestern University(统计学系,西北大学)

AI总结 提出D2AC算法,通过融合扩散策略与分布式评论家,实现无模型强化学习中扩散策略的在线高效训练,在18个困难任务上达到最先进性能。

Comments Accepted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15105 2026-05-25 cs.LG

Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting

Super-Linear: 一种轻量级预训练线性专家混合模型用于时间序列预测

Liran Nochumsohn, Raz Marshanski, Hedi Zisling, Omri Azencot

机构 * Faculty of Computer and Information Science, Ben-Gurion University(计算机与信息科学学院,本·古里安大学)

AI总结 提出Super-Linear,一种轻量级可扩展的混合专家模型,用频率特化的线性专家替代深度架构,通过光谱门控机制实现高效准确的时间序列预测。

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14311 2026-05-25 cs.LG cs.AI

Online Learning with Multiple Fairness Regularizers via Graph-Structured Feedback

通过图结构反馈进行多重公平正则化器的在线学习

Quan Zhou, Jakub Marecek, Robert Shorten

机构 * Department of Mathematics, National University of Singapore(新加坡国立大学数学系) Department of Computer Science, Czech Technical University(捷克技术大学计算机科学系) Dyson School of Design Engineering, Imperial College London(伦敦帝国理工学院设计工程戴森学院) Imperial College London(伦敦帝国理工学院)

AI总结 本文针对在线决策中多重公平约束的权重自适应问题,提出了一种基于图结构反馈的赌博机算法,能够在不预先知道权重的情况下在线学习并平衡多个公平性目标。

Comments Published in Transactions on Machine Learning Research (TMLR), 2026. OpenReview: https://openreview.net/forum?id=y8iWuDZtEw

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22098 2026-05-22 cs.CV cs.AI cs.LG

TextTeacher: What Can Language Teach About Images?

TextTeacher: 语言能教会我们关于图像什么?

Tobias Christian Nauen, Stanislav Frolov, Brian Bernhard Moser, Federico Raue, Ahmed Anwar, Andreas Dengel

机构 * RPTU University Kaiserslautern-Landau(赖兴海大学凯撒斯劳滕-兰道分校) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))

AI总结 该研究提出TextTeacher方法,通过将语言模型的语义知识注入到图像分类训练中,提升视觉模型的性能,同时保持推理时的模型简洁性。

Comments Published at TMLR

Journal ref Transactions on Machine Learning Research, ISSN 2835-8856, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03454 2026-05-22 cs.LG

[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL

[Re] FairDICE:多目标离线RL中的公平权衡

Peter Adema, Karim Galliamov, Aleksey Evstratovskiy, Ross Geurts

机构 * University of Amsterdam(阿姆斯特丹大学)

AI总结 该研究探讨了多目标离线强化学习中公平权衡的问题,提出FairDICE算法通过自适应学习多目标权重来实现公平妥协,但发现代码错误导致其在连续环境中退化为标准行为克隆,并需修正超参数以提升实验有效性。

Comments 12 pages, 8 figures in main text. Code at https://github.com/p-adema/re-fairdice. Reviewed at https://openreview.net/forum?id=Tr6MBt0hAj

Journal ref Published 05/2026 in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04371 2026-05-22 cs.AI

Cumulative Reasoning with Large Language Models

基于大语言模型的累积推理

Yifan Zhang, Jingqin Yang, Yang Yuan, Andrew Chi-Chih Yao

机构 * IIIS, Tsinghua University(清华大学人工智能研究院) Shanghai Qi Zhi Institute(上海启智研究院)

AI总结 本文提出了一种名为累积推理(CR)的框架,通过模拟人类的迭代和累积思维过程,增强大语言模型(LLM)的问题解决能力。CR通过分解任务、生成并验证中间推理步骤,构建动态有向无环图(DAG)来组成解决方案,从而在逻辑推理、24点游戏和数学问题等任务中取得了显著的性能提升。

Comments Published in Transactions on Machine Learning Research (TMLR). Project Page: https://github.com/iiis-ai/cumulative-reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04938 2026-05-22 cs.CV cs.AI cs.LG

Improved DDIM Sampling with Moment Matching Gaussian Mixtures

改进的DDIM采样与矩匹配高斯混合模型

Prasad Gabbur

机构 * Independent Researcher(独立研究者) Apple(苹果公司)

AI总结 本文提出在DDIM框架中使用高斯混合模型作为反向转换操作符,通过约束GMM参数匹配DDPM前向边缘的矩,从而在少量采样步骤下提升生成样本质量,实验表明GMM核在FID和IS指标上优于传统高斯核。

Comments 34 pages, 12 figures; Accepted to TMLR; Code open sourced

Journal ref Transactions on Machine Learning Research, 05/2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.12906 2026-05-21 cs.LG stat.ML

The Score-Difference Flow for Implicit Generative Modeling

隐式生成建模的分数差流

Romann M. Weber

机构 * Disney Research(迪士尼研究)

AI总结 本文提出分数差流作为隐式生成建模的一种新方法,通过最优减少两个分布之间的KL散度,展示了其与去噪扩散模型的等价性,并揭示了生成对抗网络训练中隐含的数据优化子问题与分数差流之间的联系。

Comments 25 pages, 5 figures, 4 tables. Updated final version of a paper originally published in Transactions on Machine Learning Research (TMLR), including minor typographical corrections and post-publication commentary connecting the SD flow to drifting models

Journal ref Transactions on Machine Learning Research (7/2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01152 2026-05-20 cs.LG cs.AI cs.CV

Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution

开放集域适应在背景分布偏移下的挑战:挑战与一种可证明高效的解决方案

Shravan Chaudhari, Yoav Wald, Suchi Saria

机构 * Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系) Faculty of Data and Decision Sciences, Technion(技术学院数据与决策科学学院) Center for Data Science, New York University(纽约大学数据科学中心) Bayesian Health(贝叶斯健康)

AI总结 本文研究了在背景分布偏移情况下开放集域适应的挑战,并提出了一种可证明高效的解决方案CoLOR,通过理论分析和实验证明其在简化过参数化设置中优于基线方法,同时展示了其在图像和文本数据上的广泛适用性。

Comments Project page at https://github.com/Shra1-25/CoLOR

Journal ref Transactions on Machine Learning Research (TMLR) 2026/May ISSN: 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05317 2026-05-20 cs.CV

ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation

ProJo4D:渐进式联合优化用于稀疏视图逆物理估计

Daniel Rho, Jun Myeong Choi, Biswadip Dey, Roni Sengupta

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Meta Reality Labs(Meta现实实验室)

AI总结 本文提出ProJo4D,一种渐进式联合优化框架,用于解决稀疏视图下逆物理参数估计问题,通过逐步扩展联合优化参数集,提高了4D未来状态预测和物理参数估计的准确性,达到几何精度提升10倍的性能。

Comments TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08248 2026-05-20 cs.CV

TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation

TextBoost: 通过文本编码器提升文本到图像生成的个性化

NaHyeon Park, Kunhee Kim, Hyunjung Shim

机构 * KAIST(韩国科学技术院)

AI总结 本文提出TextBoost,一种高效的文本到图像扩散模型单次个性化方法,通过仅微调文本编码器提升计算和存储效率,并保持语义完整性,从而实现更快收敛和更低存储需求,同时保持高质量生成。

Comments Project page: https://textboost.github.io. Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18534 2026-05-19 cs.LG

XCTFormer: Leveraging Cross-Channel and Cross-Time Dependencies for Enhanced Time-Series Analysis

XCTFormer: 利用跨通道和跨时间依赖性提升时间序列分析

Israel Zexer, Omri Azencot

机构 * The Stein Faculty of Computer and Information Science(施坦计算机与信息科学系) Ben-Gurion University of the Negev(本·古里安大学)

AI总结 本文提出XCTFormer模型,通过增强的注意力机制显式捕捉时间序列中的跨时间与跨通道依赖性,以提升时间序列分析性能,特别是在缺失值填补任务中取得state-of-the-art结果。

Comments TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08080 2026-05-19 cs.LG cs.NE stat.AP

Symbolic Quantile Regression for the Interpretable Prediction of Conditional Quantiles

符号量化回归用于条件量化可解释性预测

Cas Oude Hoekstra, Floris den Hengst

机构 * Independent researcher(独立研究者) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

AI总结 本文提出了一种符号量化回归方法,用于预测条件量化并解释预测变量对结果的影响,通过在航空燃料使用案例中比较预测极值和中央结果的模型,展示了SQR在高风险应用中的有效性。

Journal ref Transactions on Machine Learning Research, May 2026, https://openreview.net/pdf?id=x9OYbyPJOG

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24497 2026-05-19 cs.AI cs.LG cs.RO stat.ML

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

在联合嵌入预测世界模型中成功因素是什么?

Basile Terver, Tsung-Yen Yang, Jean Ponce, Adrien Bardes, Yann LeCun

机构 * Meta FAIR Inria Paris(巴黎理工院) Ecole normale supérieure / PSL(巴黎高等师范学院 / PSL) New York University(纽约大学)

AI总结 本文研究了在物理规划中使用联合嵌入预测世界模型(JEPA-WMs)的成功因素,通过分析模型架构、训练目标和规划算法对规划成功的影响,提出了一种在导航和操作任务中优于现有基线方法的模型。

Comments V2 of the article: - Added AdaLN-zero - Added table comparing JEPA-WMs with baselines with std translating per-seed variability only, no variability across epochs - Reordered figures in main body of the paper V3: added data scaling experiments, theoretical appendix section on autoregressive rollout, acceptance at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13846 2026-05-19 cs.CL cs.AI cs.LG

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

LightTransfer: 你的长上下文LLM实际上是一个具有轻松适应能力的混合模型

Xuan Zhang, Fengzhuo Zhang, Cunxiao Du, Chao Du, Tianyu Pang, Wei Gao, Min Lin

机构 * Singapore Management University(新加坡国立大学) National University of Singapore(新加坡国立大学) Sea AI Lab, Singapore(新加坡海智实验室)

AI总结 本文提出LightTransfer方法,通过将LLaMA等模型转换为混合架构,实现更高效的生成,实验表明在长上下文理解任务中,即使有半数层被识别为懒层,也能在性能损失小于1.5%的情况下提升2.17倍的吞吐量,并在数学基准AIME24上达到53.3%的分数。

Comments Accepted by TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16318 2026-05-19 cs.LG

Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning

在强化学习中探究循环神经网络的动作编码

Matthew Schlegel, Volodymyr Tkachuk, Adam White, Martha White

机构 * University of Alberta(阿尔伯塔大学)

AI总结 本文探讨了在强化学习中如何通过修改循环神经网络架构来整合动作信息,评估了不同方法在多个示例领域中的效果,并讨论了未来发展的挑战。

Comments Published in TMLR in 2023, https: // openreview. net/ forum? id= K6g4MbAC1r .Transactions on Machine Learning Research (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18225 2026-05-18 cs.LG stat.ML stat.OT

Adaptive Conformal Prediction for Quantum Machine Learning

适应性符合预测用于量子机器学习

Douglas Spencer, Samual Nicholls, Michele Caprio

机构 * Mathematical Institute, University of Oxford(牛津大学数学研究所) Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系)

AI总结 本文提出适应性量子符合预测算法,解决量子处理器时间变化噪声对符合保证的影响,通过重复校准保持有效性,实验证明其在IBM量子处理器上的稳定性和覆盖率。

Comments Accepted at TMLR 05/2026. 27 pages, 5 figures

Journal ref Transactions on Machine Learning Research, May 2026, ISSN 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15484 2026-05-18 cs.CV cs.LG

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing

何时稀疏MoE在视觉中起作用?背骨计算利用在稀疏路由中的作用

Libo Sun, Po-wei Harn, Peixiong He, Xiao Qin

机构 * Department of Computer Science and Software Engineering(计算机科学与软件工程系) Auburn University(阿伯拉罕大学) Department of Information Management(信息管理系) National Central University(国立中央大学)

AI总结 研究稀疏top-k路由在视觉分类中的有效性,发现计算利用模式,指出背骨架构和多专家路由对性能的影响,通过实验验证关键因素。

Comments 24 pages (main + appendix), 8 figures, 18 tables. Under review at TMLR. Code and aggregate results: https://github.com/libophd/sparse-moe-vision-rho

详情

展开后加载摘要…

URL PDF HTML 收藏