Cloudflare Clef models reach security ops as Musubi ships PolicyLM
Cloudflare released Apache 2.0 Clef decision models, later using Clef to filter security alerts, while Musubi shipped PolicyLM-1.7B for moderation.
Cloudflare's Workers AI team released Clef and Clef-flash, Apache 2.0 decision models post-trained from Qwen3.8-27B and Qwen3.5-9B that emit probabilities for typed questions instead of generated text. On Cloudflare's own Decision Index 0.2.1, a Clef model scored highest on 7 of 10 benchmarks, with BANKING77 macro-F1 of 94.20 against Jev's 79.74, while Jev leads GPQA Diamond, MMLU-Pro, and BBH; all of those figures are vendor-reported. MarkTechPost gives median latencies of 209.3 ms and 38.8 ms, which The Decoder rounds to about 209 ms and 39 ms, versus just over 524 ms for TypeSafe AI's Jev (spelled Typesafe in the Musubi report). The models offer a 64,000-token context, image input, Workers AI and Hugging Face availability, and fine-tuning via reinforcement learning with Replicate technology; The Decoder also reports Cloudflare saying agents need no human in the loop. Separately, Musubi released open-weights PolicyLM-1.7B for real-time moderation under 50 ms, returning a binary category match without retraining when policy text changes. Cloudflare's later Managed Defense write-up is narrower than that no-human claim: deterministic evidence is collected first, Clef filters likely false positives, and remaining alerts go to constrained specialists plus approved models such as GPT-5.6 Cyber and Mythos because a single-agent prototype had hallucinated unsupported claims.
- On 2026-10-02 Cloudflare released Clef (27B, post-trained from Qwen3.8-27B) and Clef-flash (9B, from Qwen3.5-9B) as Apache 2.0 open-weight models that return probabilities for typed questions rather than free-form text, retaining vision…
- On vendor-reported Decision Index 0.2.1, a Clef model led 7 of 10 benchmarks, including BANKING77 macro-F1 94.20 versus Jev 79.74; Jev still leads GPQA Diamond, MMLU-Pro, and BBH, with no independent replication stated.
- Median latency is 209.3 ms for Clef and 38.8 ms for Clef-flash (rounded to about 209 ms and 39 ms in a second account), versus just over 524 ms for TypeSafe AI's Jev.
- Both support a 64,000-token context and images, are on Hugging Face and Workers AI, and can be fine-tuned through a reinforcement-learning service using Replicate technology.
- Musubi's PolicyLM-1.7B, announced 2026-10-06, applies plain-English moderation policies in under 50 ms, returns a binary judgment, and does not need retraining when the policy changes; Musubi links it to 2024 GLiNER and compares it with…
- Cloudflare's 2026-10-07 Managed Defense harness collects a fixed evidence snapshot before inference, uses Clef to score likely false positives, then specialist agents and approved models including GPT-5.6 Cyber and Anthropic's Mythos,…
Coverage timelineoldest first · each row is one article
- · 6d agoCloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
MarkTechPost· 60
Cloudflare released Clef (27B) and Clef-flash (9B), Apache 2.0 decision models that return typed probabilities.
- · 6d agoCloudflare says its new Clef model means humans no longer need to be in the loop for AI agents
The Decoder· 58
Cloudflare launched open-weight Clef and Clef-flash decision models so AI agents can act without a human in the loop.
- · 2d agoHow AI decision models could change content moderation
TechCrunch · AI· 44