arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27648cs.SE

换个名字就不甜了:LLM文档工作流中格式鲁棒性的实证研究

Not as Sweet by Another Name: An Empirical Study of Format Robustness in LLM Document Workflows

Xiaoyu Zhang, Xianyun Cheng, Tianlin Li, Yuwei Zheng, Yue Yang, Yang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出感知格式的变形测试框架,经大规模实证研究发现LLM文档工作流的格式鲁棒性存在严重问题,切换格式可致准确率降53.63%,并设计出无需重训模型的缓解策略。

中文摘要 AI 辅助

由大语言模型(LLM)驱动的软件系统正快速从纯文本对话向以文档为中心的端到端工作流演进,相同语义内容可通过文件上传接口以多种文档格式(如CSV)交付。但现有测试工作聚焦于输入为单一提示字符串的模型与系统的鲁棒性和可靠性,留下关键问题未解答:当相同内容以不同文档格式传入时,这些文档工作流能否保持鲁棒行为?为填补该空白,本文提出一种感知格式的变形测试框架,包含三条变形关系,以全面评估端到端LLM文档工作流的格式鲁棒性。基于该框架,我们开展了大规模实证研究,覆盖4个代表性LLM工作流、4个真实世界任务和4种文档格式,共包含48000次工作流执行。研究发现,格式变化会带来系统性且严重的威胁:仅切换格式就可能导致准确率下降最多53.63%,并在超过41%的实例中触发决策漂移。我们进一步从用户视角设计轻量级缓解策略,无需模型重新训练即可恢复多达44.21%的格式诱导决策漂移。本研究表明,文档格式并非中性包装,而是影响LLM软件系统可靠性的关键因素,呼吁在高风险场景实际部署时采用相应的测试与保障措施。

英文摘要

LLM-driven software systems are rapidly evolving from plain-text conversations to document-centric end-to-end workflows, where the same semantic content can be delivered in diverse document formats (e.g., CSV) through file upload interfaces. Yet existing testing work focuses on the robustness and reliability of models and systems whose input is a single prompt string, leaving a critical question unanswered: Can these document workflows maintain robust behaviors when the same content arrives in a different document format? To fill the gap, in this paper, we propose a format-aware metamorphic testing framework with three metamorphic relations to comprehensively evaluate the format robustness of end-to-end LLM document workflows. Based on this framework, we conduct a large-scale empirical study spanning four representative LLM workflows, four real-world tasks, and four document formats, comprising a total of 48,000 workflow executions. Our findings reveal that format variation poses a systematic and serious threat. Merely switching formats can cause accuracy to drop by up to 53.63% and trigger decision drifts in over 41% of instances. We further design lightweight mitigation strategies from the users' perspective that recover up to 44.21% of format-induced decision drift without model retraining. Our study demonstrates that document format is not a neutral wrapper but a critical factor affecting the reliability of LLM software systems, calling for corresponding testing and safeguards in the deployment in real-world high-stakes scenarios.

补充信息

↑