Adding AI to a Security Toolkit: Start With Your Own Scripts
Open your shell history before you open a course catalog. The jq filters, the grep -v chains against Zeek logs, the PowerShell one-liners you paste into a ticket every week: that is your toolkit. Adding AI to it means replacing one step in one of those pipelines with something that does the step better. It does not mean learning "AI" as a separate subject and hoping it attaches to your job later.…
Before delving into the world of AI, start with the tools you already have. Examine your own scripts, such as jq filters, grep -v chains against Zeek logs, and PowerShell one-liners. These are the building blocks of your toolkit. When you incorporate AI, you're essentially enhancing one step within one of these existing pipelines.
It's not about becoming an expert in AI; instead, it's about leveraging the technology to improve what you're already doing. Many practitioners make the mistake of first learning AI as a standalone subject and then trying to integrate it into their job, but this approach rarely works. Instead, work backwards from the pipeline you're using.
There are three common pipelines most security teams run, along with the AI step that can enhance each one, and what training is required for that step. Pipeline one focuses on adjusting the threshold in your hunt script, which is often a fixed value that may not be suitable for all hosts. Instead of relying on a single threshold, compute a per-host baseline using pandas and calculate a robust z-score.
This will enable you to flag any host that sends an unusual amount of data outbound, regardless of the host's typical traffic pattern. Pipeline two deals with obfuscated script blocks, such as PowerShell script block text that may contain base64 encoding, string reversal, and -join tricks. Decoding this manually can be time-consuming.
An LLM (large language model) can help in the initial pass of decoding the script block, and tools like llm CLI can be used to structure the output. However, it's important to remember that the script block is attacker-authored, so treat the model's output as a lead for the analyst and never close an alert based on the model's output alone.
Pipeline three involves the LLM feature that your company has recently shipped. If you already run web application tests, you have a pipeline that includes scope, enumeration, testing, and reporting. Your organization's support chatbot or internal RAG (retrieval-augmented generation) assistant can be integrated into this pipeline.
NVIDIA's garak is a good example of a scanner-style entry point that can be used to test your HTTP endpoint with a YAML configuration file. Keep in mind that the LLM output is nondeterministic, so one clean run does not prove much. To report against MITRE ATT&CK and OWASP Top 10 for LLM Applications, you'll need to train your team on mapping the application's data path, understanding manual attacks that scanners may miss, and reporting findings accordingly.
If you cannot yet read log files and understand what normal behavior looks like, AI won't be able to help you. Instead, focus on learning the data first. Similarly, if your team's volume is low, a fixed threshold on a network of a few hosts can be tuned manually without the need for an LLM.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.