评估网络安全中大语言模型生成的检测规则
文章背景与核心概要
随着大语言模型(LLM)在网络安全运营中的普及,如何建立信任并衡量其实际运营效果依然是一个巨大的挑战。本文介绍了一个开源评估框架和基准指标,旨在严格评估由大语言模型生成的安全规则。该基准采用基于留出集(holdout set)的方法,将AI生成的规则与人工编写的规则集进行对比,并结合了三个专家启发的指标,对自动化规则生成器进行现实且多维度的评估。
本文通过Sublime Security检测团队的规则以及其自动化检测工程师(ADE)生成的规则展示了该方法论,并对ADE的功能进行了详细分析。
📌 Summary
As Large Language Models (LLMs) become increasingly pervasive in cybersecurity operations, establishing trust and measuring their operational effectiveness remains a significant challenge. This paper introduces an open-source evaluation framework and benchmark metrics designed to rigorously assess LLM-generated security rules. Utilizing a holdout set-based methodology, the benchmark compares AI-generated rules against human-curated corpora. It incorporates three expert-inspired metrics to deliver a realistic, multifaceted evaluation of automated rule generators. The methodology is demonstrated using rules from Sublime Security's detection team alongside those generated by their Automated Detection Engineer (ADE), complete with a detailed analysis of ADE's capabilities.
📋 Paper Metadata
Field Details arXiv Identifier arXiv:2509.16749[cs.CR]Primary Subject Cryptography and Security ( cs.CR)Secondary Subjects Artificial Intelligence ( cs.AI)Submission Date September 20, 2025 Conference Venue Accepted at the Conference on Applied Machine Learning in Information Security (CAMLIS 2025) Document Stats 11 pages, 3 figures, 4 tables License Creative Commons Attribution 4.0
👥 Authors
- Anna Bertiger
- Bobby Filar
- Aryan Luthra
- Stefano Meschiari
- Aiden Mitchell
- Sam Scholten
- Vivek Sharath
🔗 Access & Resources
Full-Text Links
External Tools & Citations
- Citations: Google Scholar | Semantic Scholar | NASA ADS
- Code & Exploration: Hugging Face | CatalyzeX Code Finder | alphaXiv
