HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents
HESP, an external controller, lets small local LLMs complete SOC alert triage instead of stalling without a verdict.
HESP is a controller that keeps SOC alert-investigation procedure outside a local open-weight LLM. It tracks competing explanations, picks read-only probes by information gain per cost, accepts only evidence-backed verdicts, and can end an investigation itself. In 7,272 audited episodes across five models from 7B to 72B, it raises Qwen2.5-7B verified completion from 0.125 to 1.000 and lifts Llama-3.1-8B, which never concludes alone, from 0 to 0.917. Code, protocols, and episode journals are released for on-premises use.
- HESP selects read-only probes by expected information gain per cost.
- Qwen2.5-7B verified completion rises from 0.125 to 1.000.
- A controller-side stop lifts Llama-3.1-8B from 0 to 0.917.
- Study covers 7,272 episodes on models from 7B to 72B.
Full article233 words · extracted from arxiv.org · click to collapse
Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and small local models fail at it: they probe without converging, never commit to a verdict, or dismiss real attacks. In this paper, we present HESP, a controller that holds the investigation procedure outside the model. HESP keeps a ledger of competing explanations, selects read-only probes by expected information gain per cost, accepts only verdicts backed by current evidence, can end an investigation itself, and journals every prediction before its observation. We evaluated HESP in four pre-registered studies with five open-weight models from two families (7B to 72B), totalling 7,272 audited episodes in a controlled triage environment. With likelihood tables counted from LLM-free runs, HESP lifts Qwen2.5-7B from 0.125 to 1.000 verified completion, matching oracle tables. The information-gain ranking adds +0.26 to +0.35 on every model that concludes, and a controller-side stop lifts Llama-3.1-8B, which never concludes on its own, from 0 to 0.917. What to probe and when to stop are therefore separate failures, and different small models exhibit different ones. Because HESP and its planner run entirely on local hardware, it suits environments where telemetry cannot leave the premises. We release all code, protocols, and episode journals at https://github.com/lzwhehe/HESP.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33446