扩展FunctionGemma以实现实用的端侧移动设备函数调用
Extending FunctionGemma for Practical On-Device Mobile Function Calling
- UGrowAI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究扩展FunctionGemma 270M-it,通过构建约9,500个对话的合成数据集MOBILEACTIONSEXTENDED并微调模型,将端到端准确率提升至76.5%,同时保持紧凑模型,为低延迟、隐私保护的端侧移动助手提供实用方案。
AI中文摘要:
端侧助手需要能够将自然语言映射到本地系统操作的函数调用模型,但现有资源侧重于Web API或狭窄的移动操作目录。我们通过引入MOBILEACTIONSEXTENDED来扩展FunctionGemma 270M-it,以适用于实际的Android工作流程。MOBILEACTIONSEXTENDED是一个合成的、经过模式验证的数据集,包含约9,500个对话,覆盖十五个设备控制类别,包括消息传递、电话呼叫、相机/截图、亮度控制、设备状态查询、手电筒控制以及应用程序管理。我们使用TRL监督微调,在仅完成损失下对270M模型进行微调,生成了一个扩展专家模型和一个与Google的MOBILEACTIONSGOOGLE联合训练的合并模型。在MOBILEACTIONSEXTENDED上,端到端准确率从基础模型的29.3%和Google的Mobile-Actions变体的17.2%提升至76.5%。合并模型在MOBILEACTIONSEXTENDED上保持76.5%的准确率,并在MOBILEACTIONSGOOGLE上达到82.3%,相比Google的Mobile-Actions专家模型的90.3%有所下降,这代表了8.0个百分点的权衡,以换取类别覆盖范围翻倍。我们发布了数据集、微调模型、可复现的训练/评估流程以及一个Android演示,强调了紧凑的本地函数调用作为实现低延迟和隐私保护的移动助手的实用途径。
英文摘要:
On-device assistants require function-calling models that map natural language to local system actions, but existing resources emphasize web APIs or narrow mobile-action catalogs. We extend FunctionGemma 270M-it to practical Android workflows by introducing MOBILEACTIONSEXTENDED, a synthetic, schema-validated dataset of ~9,500 conversations covering fifteen device-control categories, including messaging, phone calls, camera/screenshot, brightness control, device-status queries, flashlight control, and application management. We fine-tune the 270M model with TRL supervised fine-tuning under completion-only loss, producing an extended specialist and a combined model trained jointly with Google's MOBILEACTIONSGOOGLE. On MOBILEACTIONSEXTENDED, end-to-end accuracy improves from 29.3% for the base model and 17.2% for Google's Mobile-Actions variant to 76.5%. The combined model retains 76.5% on MOBILEACTIONSEXTENDED and reaches 82.3% on MOBILEACTIONSGOOGLE, down from the 90.3% of Google's Mobile-Actions specialist, representing an 8.0-percentage-point trade-off in return for doubling category coverage. We release the dataset, fine-tuned models, reproducible training/evaluation pipeline, and an Android demo, highlighting compact local function calling as a practical path towards low-latency and privacy-preserving mobile assistants.