Residual Community Prototypes Under-Reject Held-Out Malware Families in FCG-MFD
Study finds Louvain community features fail to stably reject held-out malware families in FCG-based open-set recognition versus simple uncertainty.
The paper tests whether residualized Louvain-community summaries add rejection capability beyond a GIN embedding and dimension-matched generic topology on a deduplicated, conflict-audited FCG-MFD corpus with five held-out families. Community features did not produce stable held-out-family rejection: ranking effects reversed across families, false-positive rate at 95% unknown recall worsened for every family, and a validation-fitted threshold rejected only 4.48% of unknown samples. Accepted-known macro F1 improved in every family but the exact two-sided sign-flip p-value was 0.0625. Simple classifier uncertainty performed better on ranking, high-recall rejection, and OSCR, motivating matched topology controls and held-out-family analysis in graph open-set evaluations.
- Residual community prototypes fail to create a stable unknown margin for held-out families.
- Validation-fitted threshold rejects only 4.48% of unknown samples.
- Known-family macro F1 gains are statistically inconclusive (p=0.0625, n=5).
- Simple classifier uncertainty outperforms on ranking, high-recall rejection, and OSCR.
Full article171 words · extracted from arxiv.org · click to collapse
Open-set malware-family recognition must classify known families while rejecting families absent from training. We test whether Louvain-community summaries add rejection information beyond a graph neural network embedding and dimension-matched generic topology. The study uses a deduplicated, conflict-audited FCG-MFD corpus, five held-out families, and three optimization seeds. Community features are residualized against generic topology using known-family training data before nearest-prototype scoring. Residual community does not produce stable held-out-family rejection. Ranking effects reverse across families, the false-positive rate at 95 percent unknown recall worsens for every held-out family, and a validation-fitted threshold rejects only 4.48 percent of unknown samples. Accepted-known macro F1 improves in every family, but with five independent family units the exact two-sided sign-flip p-value is 0.0625, the smallest attainable value. The score remains associated with graph scale, while simple classifier uncertainty performs better on ranking, high-recall rejection, and OSCR. In this GIN/FCG-MFD setting, community-enriched prototypes change known-class geometry without creating a stable unknown margin. Graph open-set evaluations should pair structural features with matched topology controls, operational thresholds, and held-out-family analysis.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.24980