DS@GT ARC参与Touché任务:用于检索增强辩论的大型语言模型
DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate
浏览论文内容
中文总结 AI 辅助
DS@GT ARC参与Touché 2025检索增强辩论任务,采用三家提供商的六个主流LLMs通过检索增强提示流水线,发现多LLM评估者一致性无法可靠追踪官方评估目标,在质准则上差距最大。
中文摘要 AI 辅助
我们将DS@GT ARC的工作笔记提交扩展至Touché 2025检索增强辩论任务。该任务包含两个子任务:在模拟辩论中生成下一轮发言,以及根据Gricean(格赖斯)的量、质、关系、方式准则评估辩论回应。DS@GT ARC的提交内容由三家提供商的六个主流大型语言模型(LLMs)通过检索增强提示流水线构成。我们总结了工作论文的结果,并探究多LLM评估者的一致性是否可作为官方评估性能的可靠替代指标。分析表明,前沿LLM系统是强大的回应生成器,作为评估者时在模型家族内部一致性较强,但这种共识无法可靠追踪官方评估目标,在质准则上存在最大差距。本文配套源代码位于两个指定的https URL。
英文摘要
We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at https://github.com/dsgt-arc/touche-2025-rad and https://github.com/dsgt-arc/touche-2025-rad-analysis.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。