AI 中文总结
本文通过提示设计和微调改进本地LLM SSH蜜罐,发现提示设计影响显著,扩展数据集微调可提升性能,但提示与微调效果无法简单结合,二者可能因重叠约束产生冲突。
AI 中文摘要
基于大语言模型(LLM)的SSH蜜罐常使用闭源云端LLM,因为这类模型能提供强的Shell真实性,但云端模型会带来部署问题,包括无稳定版本控制、服务提供方变更、攻击者驱动的成本以及模型退役。本地开放权重模型可避免这些问题,但通常性能更差,且会出现暴露蜜罐的错误,这些错误包括畸形输出、命令回显、不一致的文件系统状态以及AI风格的人工制品。本文研究如何通过提示设计和监督微调来改进和评估基于本地LLM的SSH蜜罐的Shell仿真准确性。我们总共对8个模型进行了微调与评估:shelLM中使用的原始微调GPT-3.5模型,以及7个本地开放权重模型,每个模型都与其基础模型进行了比较。我们还测试了提示结构在不同模型家族间的迁移情况。使用34个自动化单元测试(用于测量单会话和新会话设置下的Shell仿真准确性),我们发现提示设计影响显著,而微调效果取决于数据集覆盖范围。对原始的112个对话数据集进行微调并未提高总通过率,而从蜜罐日志构建的扩展数据集则能生成明显更强的本地模型。综合来看,结果表明提示工程和微调各自都能独立改进本地LLM蜜罐,但它们的效果无法简单结合,因为基于强规则的提示和监督适配可能会因处理重叠的Shell行为约束而产生冲突。
英文摘要
LLM-based SSH honeypots often use closed cloud LLMs because they give strong shell realism, but cloud models create deployment problems. These include no stable versioning, provider-side changes, attacker-driven cost, and model decommissioning. Local open-weight models avoid these problems, but they usually perform worse and make mistakes that reveal the honeypot. These mistakes include malformed outputs, command echoing, inconsistent filesystem state, and AI-style artifacts. This paper studies how to improve and evaluate the shell emulation accuracy of local LLM-based SSH honeypots using prompt design and supervised fine-tuning. We fine-tune and evaluate eight models in total: the original fine-tuned GPT-3.5 model used in shelLM and seven open-weight local models, each compared to its base model. We also test how prompt structure transfers across model families. Using 34 automated unit tests that measure shell emulation accuracy in single-session and fresh-session settings, we find that prompt design has a large effect and that fine-tuning depends on dataset coverage. Fine-tuning on the original 112-conversation dataset does not improve aggregate pass rate, while an expanded dataset built from honeypot logs produces clearly stronger local models. Taken together, the results suggest that prompting and fine-tuning can each improve local LLM honeypots on their own, but their effects do not combine straightforwardly, since strong rule-based prompting and supervised adaptation can also conflict by addressing overlapping shell-behavior constraints.