ZeroHour
MarkTechPostpublished ()ingested Michal Sutter

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

infoAI tools & infraimportance 55
AI summary · glm-5.3-flash

Google open-sourced Mantis, an Apache-2.0 modular skills toolkit that lets AI coding agents find, reproduce, and patch vulnerabilities with sandboxed verification.

Google released Mantis on GitHub under Apache 2.0 as a stack-agnostic set of slash-command skills that chain through the full vulnerability lifecycle: mining version history, building threat models, filtering findings, reproducing bugs in gVisor or network-disabled VMs, assembling exploit chains, patching, and scoring residual risk from 1 to 10. It runs with Gemini CLI, Antigravity CLI, the Google ADK, or comparable agent frameworks, and a supervisor skill (/mantis-meta-agent) can drive the whole loop. Google says the design targets the sub-7 percent true-positive rate of naive AI code scanning, and that its hierarchical summary tree cuts token overhead by over 85 percent. The toolkit is deployable for local and internal evaluation but not yet recommended for production.

  • Modular slash-command skills chain discovery, dedupe, sandboxed reproduction, patching, and 1-10 risk scoring.
  • Differentiator is grounding: sandboxed reproduction and patch re-attack rather than model confidence alone.
  • Hierarchical summary tree reportedly cuts token overhead by over 85 percent.
  • Targets sub-7 percent true-positive rates of naive AI code scanning, per Google.
  • Apache 2.0, usable locally with Gemini CLI or ADK; not production-ready yet.
Full article565 words · extracted from marktechpost.com · click to collapse

Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills that lets an AI coding agent run the whole vulnerability lifecycle. It finds a suspected flaw, strips the false positives, reproduces the bug inside a sandbox, writes a minimal patch, re-attacks that patch, and scores the residual risk.

Mantis is not a scanner you aim at a repository and walk away from. It is a set of slash commands your existing coding agent loads, plus a strict set of rules about where that agent is allowed to execute code.

Is it deployable? Yes for local and internal evaluation, not yet for production. You can clone it today and run it with Gemini CLI, Antigravity CLI, the Google ADK, or any comparable agent framework.

The pipeline

Mantis publishes each stage as a separate skill directory, invoked as a slash command and chained sequentially. A supervisor skill, /mantis-meta-agent, can drive the whole loop in a long-lived session.

The early stages learn the target: /mantis-history mines version control history for past security fixes, /mantis-summarize writes the directory maps, /mantis-architecture builds a Markdown knowledge base, /mantis-threat-model derives trust boundaries, and /mantis-plan produces a targeted roadmap.

The middle stages find and filter: /mantis-researcher sweeps files against the plan, then /mantis-dedupe, /mantis-review and /mantis-critic collapse duplicates, apply negative rules, and drop issues that cannot occur in a release build.

The late stages prove and fix: /mantis-reproduce executes payloads in gVisor or a VM with networking disabled, /mantis-chain assembles multi-step exploit chains from individually confirmed findings, /mantis-patch applies and verifies the fix, /mantis-calibrate assigns a risk score from 1 to 10, /mantis-reflect writes learnings back for the next pass, and /mantis-report produces the human-readable review packet.

A newer skill, /mantis-advise, inverts the flow. It queries the accumulated threat model, past bug lineages and verified patch patterns before you write code, so the same class of bug does not land twice.

But why?

Most agentic security tooling stops at generating findings. Mantis is interesting because it treats the reproducer and the re-attack as the trust boundary, and because it publishes the inter-stage contracts so teams can wrap the skills in a deterministic harness instead of trusting an LLM to orchestrate shell commands.

Key Takeaways

  • Mantis is a modular skills toolkit for coding agents, not a standalone scanner or a supported Google product.
  • Its differentiator is grounding: sandboxed reproduction and patch re-attack, not model confidence.
  • A hierarchical summary tree cuts token overhead by over 85 percent, per Google.
  • Google cites sub-7 percent true-positive rates for naive AI code scanning as the problem Mantis targets.
  • Deployable locally under Apache 2.0 but not recommended yet for production.

Check out the google/mantis on GitHub, Agent Reference Guide, Cloud CISO Perspectives, and Getting started with Mantis. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

Michal Sutter

Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Text extracted automatically; images, tables and formatting may be missing. Original: https://www.marktechpost.com/2026/09/09/google-open-sources-mantis-a-modular-skills-toolkit-that-lets-coding-agents-find-reproduce-and-patch-vulnerabilities/