ViFA-Council:面向越南民间艺术生成的多智能体大语言模型审议框架
ViFA-Council: Multi-Agent LLM Deliberation for Vietnamese Folk Art Generation
浏览论文内容
中文总结 AI 辅助
提出ViFA-Council,一个三阶段多智能体框架,利用GPT-4o、Gemini 3.1 Pro和Claude Sonnet 4.6的协作审议,通过任务特定JSON模式约束扩散模型,提升越南民间绘画外扩和教育故事生成的文化保真度与叙事连贯性。
中文摘要 AI 辅助
本文提出了ViFA-Council,一个三阶段多智能体框架,利用多个大型语言模型(LLM)来解决两项具有文化复杂性的生成任务:基于传统越南民间绘画的图像外扩和教育故事生成。当前的单模型生成流水线常常因缺乏跨模型批判机制而遭遇风格幻觉和文化误表征问题。ViFA-Council通过编排GPT-4o、Gemini 3.1 Pro和Claude Sonnet 4.6之间的协作来应对这一挑战。它通过结构化智能体审议强制执行严格的文化约束。该审议由任务特定的JSON模式作为中介,有效衔接自然语言讨论与基于Banana Pro的扩散式图像合成。实验和用户研究表明,结构化多智能体审议是提升文化敏感、低资源艺术领域中文化保真度和叙事连贯性的一个有前景的方向。源代码和数据已在此HTTPS URL发布。
英文摘要
This paper presents ViFA-Council, a three-stage multi-agent framework that employs multiple large language models (LLMs) to tackle two culturally complex generative tasks: image outpainting and educational story generation based on traditional Vietnamese folk paintings. Current single-model generative pipelines frequently struggle with stylistic hallucinations and cultural misrepresentations because they lack mechanisms for cross-model critique. ViFA-Council addresses this challenge by orchestrating collaboration among GPT-4o, Gemini 3.1 Pro, and Claude Sonnet 4.6. It enforces rigorous cultural constraints through structured agent deliberation. This deliberation is mediated by task-specific JSON schemas that effectively bridge natural language discussions with diffusion-based image synthesis using Banana Pro. Experiments and a user study demonstrate that structured multi-agent deliberation is a promising direction for improving cultural fidelity and narrative coherence in culturally sensitive, low-resource artistic domains. The source code and data are released at https://github.com/DanielNguyen-05/ViFA-Council.
发表机构
- University of Science, Ho Chi Minh City(胡志明市科学大学)
- Vietnam National University, Ho Chi Minh City(胡志明市越南国家大学)
- University of Information Technology, Ho Chi Minh City(胡志明市信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。