arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15715cs.AI

信息提取的智能模型行为可控性:从固定工作流程到反思性智能体

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Lujia Zhang, Xingzhou Chen, Hongwei Feng

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型智能体用于信息提取任务时,反思和记忆等智能组件能否带来改进。通过会议论文数据集提取实验,比较固定工作流程与反思智能体变体,指定优化条件,评估过程行为,明确智能机制对系统行为的影响及如何推动智能体设计优化。

中文摘要 AI 辅助

大语言模型智能体越来越多地用于复杂信息提取任务,但诸如反思和记忆等智能组件是否能在固定大语言模型工作流程基础上带来可观察和可控的改进仍不明确。我们通过会议论文数据集提取来研究这个问题,系统需识别学术PDF中提及的数据集并生成结构化记录。我们将固定工作流程基线与反思智能体变体进行比较,并指定了一个优化智能体条件(S2),它通过更丰富的PDF工具和动态工具选择扩展相同任务。我们的评估强调过程级行为,同时将提取覆盖率和字段完整性作为次要结果指标。本文描述了智能机制何时改变系统行为、这些改变是否改善任务完成情况,以及观察到的失败模式如何在相同评估框架下推动优化智能体设计。

英文摘要

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows. We study this question through conference-paper dataset extraction, where a system must identify datasets mentioned in scholarly PDFs and produce structured records. We compare a fixed workflow baseline with reflective agent variants and specify an optimized agent condition (S2) that extends the same task with richer PDF tools and dynamic tool selection. Our evaluation emphasizes process-level behavior--including tool execution, retries, reflection, memory use, runtime, and failure recovery--while treating extraction coverage and field completeness as secondary outcome measures. The paper characterizes when agentic mechanisms change system behavior, whether these changes improve task completion, and how the observed failure modes motivate an optimized agent design under the same evaluation harness.

↑