The Fly That Stopped: Mushroom-Body-Inspired Habituation as a Reward-Free Scheduling Prior for Autonomous Penetration Testing
A reward-free, fly-inspired scheduler reduced duplicate actions in lab autonomous penetration tests without finding more vulnerabilities.
Researchers evaluate a reward-free scheduler for autonomous penetration testing, inspired by Drosophila mushroom-body habituation, combining sparse state encoding with decaying counters over URL classes and tool families. In a second confirmatory stage, 8 of 10 screened lab targets remained measurable, and the scheduler lowered duplicate-action ratios in all 6 non-tied pairs, with an exact one-sided p-value of 0.015625. The largest reduction was from 51 to 18 duplicate steps inside a 60-step budget. The authors say this supports the complete scheduler on that population, not an isolated ablation or a gain in vulnerability discovery.
- A reward-free habituation scheduler penalizes repeated tool and URL-class selections.
- It reduced duplicate-action ratios in all six non-tied confirmatory target pairs.
- The largest drop was 51 to 18 duplicate steps within a 60-step budget.
- Authors report no vulnerability-discovery gain and no isolated habituation ablation.
- A 13,498-neuron circuit study found no action selectivity from local plasticity.
Full article228 words · extracted from arxiv.org · click to collapse
Autonomous security-testing agents can spend much of a fixed action budget repeating earlier tool selections. We evaluate a reward-free scheduler inspired by mushroom-body novelty processing in Drosophila. It combines sparse state encoding with decaying habituation counters over structural URL classes and tool families. The counters penalize repeated clean or error outcomes without updating weights from scalar reward. Four matched campaigns motivated this design by exposing reward-accounting errors and tool-failure loops; reward-driven components did not improve the tested primary outcomes over the reward-free MB condition. A pre-registered pilot and two confirmatory stages then evaluated repeated (tool, URL) selections. In the second confirmatory stage, 8 of 10 screened lab targets remained measurable after two error-heavy slow-XSS exclusions. The habituation-enabled scheduler lowered duplicate-action ratios in all 6 non-tied target pairs (exact one-sided p=0.015625), with two ties; the largest reduction was 51 to 18 duplicate steps within a 60-step budget. This is evidence for the complete scheduler on the measurable budget-hold population, not an isolated habituation ablation or a vulnerability-discovery gain. A complementary study on a 13,498-neuron MaleCNS-derived circuit (501,267 synaptic edges with weight at least 5) found no action selectivity from the five tested local-plasticity approaches under a fixed readout; readout plasticity produced qualified positive results in synthetic tasks without establishing a biological-topology advantage. We report the population bounds, remaining input-integrity dependencies, and an internal AI-assisted review protocol alongside the results.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.29126