arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以正确方式提问:用于复杂任务解决中提示词生成的多智能体对话系统

Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S

arXiv 2608.01366首次发表:更新:

发表机构

Indian Institute of Science (IISc)(印度科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出多智能体对话系统PAWNI,通过三层提示框架优化提示生成前端,实验显示其可提升提示结构完整性、LLM输出质量并降低人类工作量,支持人-AI协作优化。

AI 中文摘要

大型语言模型(LLMs)是复杂智力任务的核心工具,但输出质量仍受限于用户提供的提示词。迭代多轮提示常导致上下文退化和认知回报递减。我们提出PAWNI(Prompt Architecture Wizard using Neural Intelligence,基于神经智能的提示架构向导),这是一个包含8个智能体的智能对话界面,通过由自进化知识库支撑的引导式问答对话,将非结构化查询转化为结构化提示。PAWNI不优化模型的响应,而是通过前置意图澄清来优化问题本身。我们还提出了包含18个提示元素的三层框架,分为Essential(核心)、Enhancement(增强)和Elevation(提升)三类。为评估系统行为并验证测量协议,我们开展了一项探索性被试内研究(样本量N=4),覆盖4项复杂任务,整合了32通道脑电图(EEG)、NASA-TLX工作量评分和行为指标。使用PAWNI时,参与者生成的提示在结构完整性上达到评估元素的42%至91%,在所有质量维度上对LLM输出的评分更高,且报告的工作量更低(NASA-TLX评分:39.6 vs 21.7)。每位参与者都在单轮中获得了满意的输出,而未借助该系统时需要1至12轮。尽管因样本量小效应量不稳定,但方向一致性支持以下假设:优化提示生成前端是人-AI协作的关键杠杆。

英文摘要

Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iterative multi-turn prompting often leads to context degradation and diminishing cognitive returns. We present PAWNI (Prompt Architecture Wizard using Neural Intelligence), an agentic conversational interface of eight agents that transforms unstructured queries into structured prompts through guided question-and-answer dialogue informed by a self-evolving knowledge base. Rather than optimising the model's response, PAWNI optimises the question itself by front-loading intent clarification. We also propose a three-tier framework of 18 prompt elements across Essential, Enhancement, and Elevation categories. To evaluate system behaviour and validate a measurement protocol, we conducted an exploratory within-subjects study (N=4) across four complex tasks, integrating 32-channel EEG, NASA-TLX workload, and behavioural metrics. Participants produced more structurally complete prompts with PAWNI (42% to 91% of assessed elements), rated LLM outputs higher across all quality dimensions, and reported lower workload (39.6 vs. 21.7 NASA-TLX). Every participant reached satisfactory output in a single turn, compared to 1-12 turns unaided. While effect sizes are unstable due to sample size, direction consistency supports the hypothesis that optimising prompt formulation front-end is a critical lever for human-AI collaboration.

Comments53 pages, 31 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑