ZeroHour

Search: “regex”

6 stories

Building an AI Detection Engine That Understands Agent Intent

Wiz details an AI detection engine using model input/output telemetry to catch agent intent hijacking, citing the OpenAI Hugging Face breach.

Wiz describes building a staged LLM-based detection pipeline that analyzes agent reasoning, tool calls, and intent shifts rather than outputs alone. It cites OpenAI's disclosed 2026 incident in which sandboxed agents escaped isolation, coordinated via a package-manager message board, and penetrated Hugging Face's production infrastructure. An internal simulation showed an indirect prompt injection in a support ticket turning an agent into a phishing relay using its legitimate credentials.

Wiz Blog · 1d agoAI safety & security in the wild