发表机构
Glassbox AI(Glassbox AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出损失引导多专家GAN框架,通过联合损失机制稳定训练,在自定义手语数据集上实现高PSNR,可部署于消费级硬件,用于手语视频合成以助力听障人士沟通。
AI 中文摘要
本初步技术报告提出一种用于手语视频合成的框架,旨在提升听障人士的沟通能力。该框架采用损失引导多专家生成对抗网络(GAN),包含三个专门的判别器——全局判别器、手部判别器和头部判别器,每个判别器分别引导生成器中对应的专家分支关注不同的视觉区域,无需显式多样性损失即可实现隐式特征专业化。为稳定该多判别器系统(其早期训练阶段会呈现混沌动态),我们引入联合损失(United Loss)共识机制,以10%的权重将每个判别器正则化至整体平均水平。各分支进一步采用双路径卷积-Transformer设计,搭配可学习的自适应特征融合(AdaptiveFeatureFusion),平衡卷积的稳定性与窗口自注意力的细节性。生成器采用三模式交替训练计划(判别器训练、整体生成、分支专用生成)。在自定义的156GB数据集(含过滤后的测试集,已移除简单和重复样本)上,我们的0.2B参数变体取得29.8的峰值信噪比(PSNR),1.3B参数变体取得30.7的PSNR,推理时的显存占用分别为1.5GB和8GB,可在消费级硬件上部署。由于单GPU训练周期需2-3个月,完整的 ablation 研究仍在进行中。该系统已在2025年香港前沿技术峰会上展示。
英文摘要
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators--global, hand, and head--each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity losses. To stabilize this multi-discriminator system, whose early-phase training otherwise exhibits chaotic dynamics, we introduce a United Loss consensus mechanism that regularizes each discriminator toward the ensemble average at a 10% weight. Each branch further adopts a dual-pathway convolutional-transformer design with learnable AdaptiveFeatureFusion, balancing the stability of convolutions against the detail of windowed self-attention. The generator is trained using an alternating three-mode schedule (discriminator, holistic generation, branch-specialized generation). On a custom 156GB dataset with a filtered evaluation set that removes easy and repetitive samples, our 0.2B-parameter variant achieves 29.78 PSNR (0.9593 SSIM), the 0.66B variant reaches 30.52 PSNR (0.9631 SSIM) after 7.37M steps, and the 1.3B variant achieves 30.72 PSNR (0.9650 SSIM). The gain from 0.66B to 1.3B is only +0.20 PSNR despite nearly doubling the parameters, demonstrating sharply diminishing returns. Inference VRAM footprints are 1.5 GB, ~5 GB, and 8 GB respectively, enabling deployment on consumer-grade hardware. Full ablation studies remain ongoing due to the 2-3 month training cycle on a single GPU. The system was showcased at the 2025 Hong Kong Frontier Technology Summit.
CommentsPreliminary technical report. 19 pages, 8 figures, 4 algorithms