发表机构
University of Notre Dame(圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FairCompressAgent智能体框架,通过语言模型规划器与执行层协作,在FPGA部署中实现公平感知压缩配置的自动选择与交互优化,显著降低存储并改善公平性。
AI 中文摘要
公平感知的模型压缩需要选择能够平衡准确性、公平性和部署成本的方法与配置。当压缩方法被组合使用或用户需求发生变化时,这些决策变得更加困难。在本文中,我们提出了FairCompressAgent(FCA),一个通过通用算子接口集成公平感知剪枝、增量量化和稀疏低秩分解的智能体框架。语言模型规划器利用模型配置文件和实测结果来选择压缩配置,而执行层则执行压缩、微调、评估和基于约束的选择。FCA还支持需求更新,并在请求无法满足时报告剩余的违规情况。在Fitzpatrick-17k数据集上使用VGG-11进行的实验比较了四种搜索方法,共涉及40种实测配置。在准确性约束请求下,FCA选择的压缩模型推理张量存储减少了59.54%,同时验证平均精度从0.5141提升至0.5233,均等机会(EOpp)从0.2251降至0.2168。在各自的停止策略下,FCA达到与一次性规划相同的最终选择,平均评估候选数分别为7.33次和12次。重复微调、留出测试和在线需求更新表征了该压缩工作流的稳定性和交互式使用。结果表明,实测反馈和显式约束支持公平感知压缩配置的选择和交互式优化。
英文摘要
Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult when compression methods are composed or the user's requirements change. In this paper, we propose FairCompressAgent (FCA), an agentic framework that integrates fairness-aware pruning, incremental quantization, and sparse low-rank factorization through a common operator interface. A language-model planner uses model profiles and measured outcomes to select compression configurations, while an execution layer performs compression, fine-tuning, evaluation, and constraint-based selection. FCA also supports requirement updates and reports the remaining violation when a request cannot be satisfied. Experiments on Fitzpatrick-17k with VGG-11 compare four search methods over 40 measured configurations. Under the accuracy-constrained request, FCA selects a compressed model with 59.54% less inference tensor storage, while validation average precision increases from 0.5141 to 0.5233 and equalized opportunity (EOpp) decreases from 0.2251 to 0.2168. It reaches the same final selection as one-shot planning with 7.33 versus 12 candidate evaluations on average, under their respective stopping policies. Repeated fine-tuning, held-out testing, and online requirement updates characterize the stability and interactive use of this compression workflow. The results demonstrate how measured feedback and explicit constraints support the selection and interactive refinement of fairness-aware compression configurations.