arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22566cs.CL

从诊断到重新设计:使用定量人种志改进多智能体大语言模型推理

From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning

  • University of California, Irvine(加利福尼亚大学欧文分校)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • Monash University(莫纳什大学)
  • Columbia University(哥伦比亚大学)
  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Vedant Khatri, Anthony Cusimano, Zachari Swiecki, Zhen Xu, Xiner Liu, Renzhe Yu

AI总结:

本研究提出定量人种志方法,以自动作文评分为例,通过分析五智能体辩论系统的交互模式诊断问题并修改提示词,使多智能体LLM的评分准确率从27.78%提升至40.28%。

AI中文摘要:

多智能体大语言模型(LLM)系统旨在通过将任务分解给具有专门功能的多个智能体来提升推理能力,但多个智能体的存在并不必然保证推理的连贯性或与任务目标一致的输出。本文提出一种定量人种志(QE)方法,用于基于智能体交互产生的对话来诊断和重新设计多智能体LLM系统。我们以自动作文评分为例测试该方法,应用认知网络分析(ENA)对一个五智能体多智能体辩论系统进行建模,考察产生正确与错误评分决策的辩论之间的差异。结果显示,在初始系统中,正确评分决策的特征是基于评分标准的论证、达成一致和阐述;相比之下,错误评分决策的特征是扩展的“命题-挑战-回应”交互,且与评分标准的关联一致性较低。随后我们利用这些发现修改智能体的提示词,修改后的系统将精确评分准确率从27.78%提升至40.28%,并使错误辩论的对话向正确辩论的基于评分标准的模式转变,让两者几乎无法区分。基于这些结果,我们认为QE可通过追踪智能体交互模式与系统性能的关联、为提示词重新设计提供依据、评估这些重新设计是否同时改变结果和交互模式,来支持AI推理的“诊断-重新设计”循环。

英文摘要:

Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent reasoning or outputs that align with task objectives. This paper introduces a quantitative ethnographic (QE) approach for diagnosing and redesigning multi-agent LLM systems based on the discourse produced through agent interactions. We test this approach using automated essay scoring as an example context, applying Epistemic Network Analysis (ENA) to model a five-agent multi-agent debate system and examine differences between debates that produced correct versus incorrect scoring decisions. Results show that, in the initial system, correct scoring decisions were characterized by rubric-grounded justification, agreement, and elaboration. Incorrect scoring decisions, in contrast, were characterized by extended proposition-challenge-response exchanges that were less consistently tied to rubric criteria. We then used the findings to revise the agents' prompts. The revised system improved exact scoring accuracy from 27.78% to 40.28% and shifted the discourse of incorrect debates toward the rubric-grounded pattern of correct ones, making the two nearly indistinguishable. Based on these results, we argue that QE can support a diagnostic-to-redesign loop for AI reasoning by tracing how patterns of agent interaction relate to system performance, informing prompt redesign, and evaluating whether those redesigns change both outcomes and interaction patterns.

补充信息

↑