Cloudflare/clef-flash — new model trending #29 on Hugging Face
Cloudflare released Clef-Flash, a 9B multimodal decision model post-trained from Qwen3.5-9B.
Cloudflare published Clef-Flash, a 9B multimodal decision model post-trained from Qwen/Qwen3.5-9B and released under Apache-2.0. It reads a state as text, JSON, images, or video plus a schema of typed questions and returns one probability per allowed option in a single forward pass, with no free-form generation. A joint schema head scores noul, choice, and score questions, and the API matches Jev and SystemOne. On Cloudflare's Decision Index 0.2.1, Clef-Flash scored 98.8 on BFCL and 93.1 on API-Bank, ahead of the larger Clef variant on those two benchmarks.
- 9B multimodal model post-trained from Qwen3.5-9B with its vision encoder.
- Outputs one logit per option; no free-form text generation or parsing.
- Reads text, JSON, images, or video and scores all questions jointly.
- Decision Index: BFCL 98.8 and API-Bank 93.1, both best in the reported table.
- Apache-2.0 release includes a joint schema head and a Jev/SystemOne-compatible API.
Full article1,568 words · extracted from huggingface.co · click to collapse
# Clef-Flash
- **Announcement:** [Clef decision models on the Cloudflare blog](https://blog.cloudflare.com/clef-decision-models)
- **Decision Index leaderboard:** [clef-evals.workers-ai-mle.workers.dev](https://clef-evals.workers-ai-mle.workers.dev)
Clef-Flash is a 9B multimodal model that turns a state and a schema of typed
questions into decisions. It reads the state as text, JSON, images, or video, and returns a
probability for every allowed option of every question in a single forward pass. There is no
free-form text generation and no output parsing.
The Clef-Flash API is fully compatible with Jev and SystemOne.
Clef-Flash is post-trained from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B). See
[Clef](https://huggingface.co/Cloudflare/clef) for the
larger variant.
## Model
- **Backbone:** Qwen/Qwen3.5-9B with its vision encoder, stored as standard sharded safetensors.
- **Joint schema head:** a small transformer head that reads the backbone's final hidden states,
routes evidence from the state to each question, and scores all options of all questions jointly.
- **Output:** one logit per allowed option for each question. Apply a softmax per question to get
probabilities.
## Files
| File | Purpose |
|---|---|
| `model-*.safetensors`, `model.safetensors.index.json`, `config.json`, `generation_config.json` | Backbone, including the vision encoder |
| `joint_head.safetensors`, `joint_head_config.json` | Joint schema head |
| `joint_schema_model.py` | Record encoding, batching, the model, `load_release_model`, and `systemone` |
| `tokenizer.json`, `tokenizer_config.json`, `chat_template.jinja`, `processor_config.json` | Tokenizer and image/video processor |
| `LICENSE` | Apache-2.0 license |
## Usage
Tested with `torch` 2.11 and `transformers` 5.10.2 on a single H200. Image and video inputs also
need `pillow`.
```python
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef-flash")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},