ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI
ORCAGen uses RAG-guided generative AI to build validated, malware-specific deception playbooks for lightweight runtime enforcement.
ORCAGen builds malware-specific deception playbooks with retrieval-augmented generation and structured prompts, then validates them offline before runtime enforcement. A curated knowledge base of malware procedures and active-defense strategies grounds generation of proof-of-concept malware and orchestration code. The authors compare GPT-4o, GPT-5.5, Gemini 3.5 Flash, Qwen3-Coder, and Claude Sonnet 4.5 on execution success, hallucination, refinement, deception effectiveness, and overhead. Across synthesized scenarios and 150 real keylogger, stealer, and ransomware samples, GPT-5.5 needed the fewest refinements and produced no observed hallucinated APIs, while Gemini 3.5 Flash had the lowest latency and runtime overhead.
- Generates PoC malware and deception code offline, then enforces only validated logic.
- A malware-procedure knowledge base grounds RAG to limit hallucinated APIs.
- Compared GPT-4o, GPT-5.5, Gemini 3.5 Flash, Qwen3-Coder, and Claude Sonnet 4.5.
- Tested synthesized scenarios and 150 real keylogger, stealer, and ransomware samples.
- GPT-5.5 needed the fewest refinements and showed no hallucinated APIs.
Full article235 words · extracted from arxiv.org · click to collapse
Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code. A curated knowledge base (KB) of malware procedures and active defense strategies grounds the generation process, helping the LLM produce threat-specific and executable deception logic rather than generic or hallucinated outputs. The generated PoC malware provides a safe and reproducible way to test whether a deception strategy can disrupt, redirect, or suppress targeted malware behavior before the strategy is added to the runtime playbook. We evaluate ORCAGen across GPT-4o, GPT-5.5, Gemini 3.5 Flash, Qwen3-Coder, and Claude Sonnet 4.5 using execution success, hallucination rate, refinement effort, deception effectiveness, and runtime overhead. The evaluation covers synthesized malware scenarios and 150 real-world malware samples across keyloggers, information stealers, and ransomware. Across the evaluated scenarios, GPT-5.5 required the fewest refinements and produced no observed hallucinated APIs, while Gemini 3.5 Flash achieved the lowest response time and runtime overhead. The results show that RAG-guided structured prompting can support the scalable construction of malware-specific deception playbooks that remain lightweight and deterministic during runtime enforcement.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.12415