arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGthoven参加SemEval-2026任务1:一个多阶段管道进入基准测试且勉强达标

RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar

Marek Šuppa, Viktória Ondrejová, Lucia Ganajová, Gregor Karetka, Daniel Skala

arXiv 2607.13189首次发表:更新:

发表机构

Comenius University in Bratislava; Cisco Systems; Zaitra s.r.o.; NaiveNeuron(布拉迪斯拉发的夸美纽斯大学; 思科系统公司; 扎伊特拉有限公司; 天真神经元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍用于SemEval-2026任务1子任务A的RAGthoven系统,它将幽默文本生成分解为多阶段LLM管道,经多次实验优化,最终配置用RAG增强规划器,还评估两种智能变体,虽工具调用预算增加,但在英语样本上未超非智能管道,在三种语言中与基线并列第一,显示语言相关回报递减。

AI 中文摘要

我们展示了RAGthoven,这是我们用于SemEval-2026任务1(MWAHAHA)子任务A(英语、西班牙语和中文的多语言受限幽默生成)的系统。RAGthoven将创造性文本生成分解为一个基于计算幽默理论(良性违规理论、基于脚本的幽默语义理论)并经过十次实验优化的多阶段大语言模型(LLM)管道(规划器、最佳N生成器、自我批判反射器、LLM裁判)。在最终配置中,我们用来自精心策划的笑话语料库的检索增强生成(RAG)增强规划器,用不同的笑话机制进行生成。我们还评估了两种智能变体——ReAct风格的顺序工具调用(实验09)和自主多分支编排(实验10)——它们通过确定性的约束审核检查器展示相同的四个阶段。在一个保留的12实例英语样本上的四个前沿模型中,尽管工具调用预算大幅增加,但两种智能变体都没有产生我们认为优于非智能管道的输出。RAGthoven在所有三种语言中与Gemini 2.5 Flash基线并列第一,组织者报告的置信区间重叠。在西班牙语中,它比基线领先42个原始Elo点(1182对1140),而在英语(1045对1081)和中文(1045对1053)中,基线在相同的统计平局中保持较高的原始评分。这些结果共同表明,一旦有强大的前沿模型参与,精心的多阶段提示工程和智能脚手架在语言方面的回报会递减。

英文摘要

We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese). RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Reflector for self-critique, LLM-as-a-judge Judge) grounded in computational humor theory (Benign Violation Theory, Script-based Semantic Theory of Humor) and refined across ten experiments. In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, seeding generation with diverse joke mechanisms. We also evaluate two agentic variants -- ReAct-style sequential tool-calling (Exp09) and autonomous multi-branch orchestration (Exp10) -- that expose the same four stages with a deterministic ConstraintAudit checker. Across four frontier models on a held-out 12-instance English sample, neither agentic variant produced outputs we judged superior to the non-agentic pipeline despite substantially higher tool-call budgets. RAGthoven shares Rank 1 with the Gemini 2.5 Flash baseline in all three languages, with overlapping organizer-reported confidence intervals. In Spanish, it leads the baseline by 42 raw Elo points (1182 vs. 1140), while in English (1045 vs. 1081) and Chinese (1045 vs. 1053) the baseline holds the higher raw rating within the same statistical tie. Together, these results suggest language-dependent diminishing returns from elaborate multi-stage prompt engineering and agentic scaffolding once a strong frontier model is in the loop.

CommentsSemEval-2026 Task 1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑