arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14587cs.AI

一种结合规则与大语言模型(LLM)的智能体框架,用于描述性文档布局的嵌入与标注:植物科学应用案例

An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case

Nicolas Turenne, Youcef Sklab, Eric Chenin, Jean-Daniel Zucker

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出结合规则与LLM的智能体框架,用于植物性状提取,在三个区域植物数据集上实现高效标注,提升了性状覆盖度与标注量,验证了其稳健性与可扩展性。

中文摘要 AI 辅助

背景:信息检索(IR)领域的最新进展结合了稠密与稀疏表示、大语言模型(LLM)及专用检索模型,以提升排序准确性、相关性和跨语言性能。段落索引、文档布局分析、语义知识表示等互补技术通过捕捉细粒度上下文与结构信息,进一步增强了检索有效性。新兴的智能体LLM框架通过支持规划、迭代推理、工具使用及多智能体协作,拓展了在不同领域的应用,同时强调严格评估、伦理考量与可信性,确保在真实场景中负责任部署。我们提出一种模块化的智能体流水线,用于植物性状提取:光学字符识别(OCR)将PDF转换为机器可读文本,分割与索引按属和种组织内容;基于规则的解析器提取结构化植物性状,大语言模型(LLM)集成体扩展性状词汇并解决歧义。该方法确保准确的物种识别、可扩展的标注及可解释的文本植物描述整合,实现对大型植物语料库的稳健且可解释的数据提取。结果:使用三个区域植物数据集,本系统在4961个物种中提取了55737个性状标注,平均每个物种9.1个性状;基于LLM的富集整合使75%的性状覆盖度提升,总标注量增加59%;OCR引擎的选择对物种识别影响较小,整体标注数量保持稳定,证明该流水线在大规模植物性状提取中的稳健性、可扩展性与可靠性。

英文摘要

Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation further enhance retrieval effectiveness by capturing fine-grained contextual and structural information. Emerging agentic LLM frameworks extend these capabilities by enabling planning, iterative reasoning, tool use, and multi-agent collaboration, thereby broadening applications across diverse domains. These frameworks also emphasize rigorous evaluation, ethical considerations, and trustworthiness, ensuring responsible deployment in real-world settings. We propose a modular, agent-based pipeline for botanical trait extraction. Optical character recognition (OCR) converts PDFs into machine-readable text, while segmentation and indexing organize content by genus and species. Rule-based parsers extract structured botanical traits, and ensembles of large language models (LLMs) expand trait vocabularies and resolve ambiguities. This approach ensures accurate species recognition, scalable annotation, and explainable integration of textual botanical descriptions, enabling robust and interpretable data extraction across large botanical corpora. Results: Using three regional botanical datasets, our system extracted 55,737 trait annotations across 4,961 species, averaging 9.1 traits per species. Integration of LLM-based enrichment improved coverage for 75% of traits, increasing total annotations by 59%. While the choice of OCR engine had a minor effect on species recognition, overall annotation counts remained stable, demonstrating the robustness, scalability, and reliability of the pipeline for large-scale botanical trait extraction.

↑