arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

YOLO-PEFT:针对YOLO系列的参数高效微调

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Xu Lin, WenJie Nie, Jinlong Peng, Weifu Fu, YueXiao Ma, Xiawu Zheng, Yong Liu

arXiv 2608.07051首次发表:更新:

AI 中文总结

YOLO-PEFT是一种针对YOLO系列的结构感知参数高效微调框架,将适配器部署建模为约束规划问题,在VOC数据集上的实验显示其检测精度优于全微调,且能减少训练内存。

AI 中文摘要

从语言模型迁移而来的通用参数高效微调(PEFT)方法在实时检测器上可能会失效,因为实时检测器的异构算子和检测专用组件存在常规Transformer栈中没有的部署约束。我们提出YOLO-PEFT,这是一种结构感知框架,将适配器部署建模为可审计的约束规划问题。给定检测器图、PEFT请求和资源预算,YOLO-PEFT会分配算子和语义角色,评估显式的算子有效性、检测器语义、图接口和部署谓词,记录每个被排除模块的原因代码,最终输出预算内的目标模块规划,或在训练前返回弃权(不执行)。在官方VOC07+12训练集验证集到VOC07测试集的协议下,规划器选定的RS-LoRA在YOLO11s上达到0.7138的mAP50-95,在YOLO12s上达到0.7307的mAP50-95,而全微调(Full-SFT)分别为0.6428和0.6662。在RT-DETR-L上,所有7种评估的LoRA系列配置均超过预定义的灾难性阈值,支持在评估覆盖范围内做出校准后的弃权(不执行)全微调决策。受控的YOLO11审计进一步显示,LoRA将峰值训练内存减少了43.9%,但训练时间延长了1.72倍。在评估的检测器家族、部署策略和校准覆盖范围内,YOLO-PEFT用显式、可检查的规划替代了手动目标模块试错,同时保留了经过验证的训练-保存-合并-导出路径;对未见过的检测器架构的弃权(不执行)仍是一个未解决的验证问题。项目主页:this http URL

英文摘要

Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specific components impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a structure-aware framework that formulates adapter placement as an auditable constraint-planning problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface, and deployment predicates, records a reason code for each excluded module, and either emits a budgeted target-module plan or returns Refuse before training. Under the official VOC07+12 trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95 on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent, although training takes 1.72 times longer. Within the evaluated detector families, placement policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal on unseen detector architectures remains an open validation problem. Project Page: github.com/Tencent/YOLO-Master

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑