发表机构
Aegix Insight(安捷智研)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Aegix Pulse提出三阶段可追踪架构,分离任务角色、品牌画像与生成修订,实验显示品牌上下文和上下文保持修订有初步提升但未达统计显著性。
AI 中文摘要
生产级内容生成系统必须整合用户的即时任务、长期品牌标识、历史证据和修订反馈。我们提出了Aegix Pulse,一种面向生产环境的三阶段架构,它将当前任务澄清与Task Persona(任务角色)确定、长期Account Profile(账户画像,即品牌DNA)组装,以及受控生成与修订分离,同时跨内容版本保留来源信息。我们使用96个合成社交媒体生成任务评估了四项预注册声明。四个初始生成条件逐步引入了Task Persona、Account Profile和成功历史风格证据,而两个修订条件比较了普通修订与上下文保持修订。该实验产生了480条完整的生成记录和1,440次盲法LLM-Judge(大语言模型评审)评估,并辅以人工审查。与仅使用Task Persona相比,添加Account Profile使平均品牌一致性得分在五分制上提高了0.1562分(Holm校正后p=.1224)。在修订期间保留任务和品牌上下文,与普通修订相比,使平均任务保持得分提高了0.2917分(Holm校正后p=.2432)。经过多重比较校正后,这两项改进均未达到统计学上的决定性结论。仅使用Task Persona显示出较小的观察效应,而在当前设置下,成功历史证据未对品牌一致性提供额外改进。人工验证未能一致地复现LLM-Judge的效应方向,且评审者间一致性较低。这些发现为持久品牌上下文和上下文保持修订提供了初步证据,同时确定了加强证据处理和评估的优先事项。
英文摘要
Production content-generation systems must integrate a user's immediate task, long-term brand identity, historical evidence, and revision feedback. We present Aegix Pulse, a production-oriented three-stage architecture that separates current-task clarification and Task Persona finalization, long-term Account Profile (Brand DNA) assembly, and controlled generation and revision while preserving provenance across content versions. We evaluate four preregistered claims using 96 synthetic social-media generation tasks. Four initial-generation conditions progressively introduced a Task Persona, Account Profile, and successful-history style evidence, while two revision conditions compared plain and context-preserving revision. The experiment produced 480 completed generation records and 1,440 blinded LLM-Judge evaluations, supplemented by human review. Adding the Account Profile increased mean brand-consistency scores by 0.1562 points on a five-point scale compared with Task Persona alone (Holm-adjusted p=.1224). Preserving task and brand context during revision increased mean task-preservation scores by 0.2917 points compared with plain revision (Holm-adjusted p=.2432). Neither improvement was statistically conclusive after multiple-comparison correction. Task Persona alone showed a small observed effect, while successful-history evidence provided no additional improvement in brand consistency under the current setting. Human validation did not consistently reproduce the LLM-Judge effect directions and showed low inter-reviewer agreement. These findings provide preliminary evidence for persistent brand context and context-preserving revision while identifying priorities for stronger evidence processing and evaluation.
Comments20 pages, 5 figures, 3 tables