AI 中文总结
本文提出一种多语言微调与阈值校准方法,用于多元价值对齐,在三个国家数据集上取得宏平均准确率0.805。
AI 中文摘要
我们介绍了针对 PlurVA-LLM 2026 共享任务赛道1 的系统,该任务聚焦于中国、印度尼西亚和斯里兰卡背景下的多元价值对齐。针对资源受限的赛道,我们使用 4 位 QLoRA 对 Llama 3.1 8B Instruct 进行了微调。我们的方法结合了针对中文数据的选项排列增强、针对印度尼西亚数据的注释者投票扩展,以及针对斯里兰卡数据的二元重构与 SinhalaMMLU 增强。我们进一步对斯里兰卡数据的预测应用了条件阈值校准。最终系统在中文、印度尼西亚和斯里兰卡数据上分别达到了 0.785、0.715 和 0.916 的准确率,总体宏平均准确率为 0.805。
英文摘要
We present our system for the PlurVA-LLM 2026 Shared Task Track-1, which focuses on pluralistic value alignment in the contexts of China, Indonesia, and Sri Lanka. For this resource-constrained track, we fine-tuned Llama 3.1 8B Instruct using 4-bit QLoRA. Our approach combines option-permutation augmentation for Chinese data, annotator vote expansion for Indonesian data, and binary reformulation with SinhalaMMLU augmentation for Sri Lankan data. We further applied conditional threshold calibration to the predictions for the Sri Lankan data. The final system achieved accuracies of 0.785 for Chinese, 0.715 for Indonesian, and 0.916 for Sri Lankan, resulting in an overall macro-average accuracy of 0.805.
Comments11 pages, 6 figures, 6 tables, Accepted paper at the first workshop on Pluralistic Value Alignment of LLMs @ AACL-IJCNLP 2026