Cloudflare/clef — new model trending #27 on Hugging Face
Cloudflare released Clef, a 27B open multimodal model that answers schema-constrained decisions in one forward pass.
Cloudflare published Clef, a 27B multimodal decision model post-trained from Qwen3.8-27B and released under Apache-2.0. It reads text, JSON, images, or video plus a schema of typed questions and returns one probability per allowed option in a single forward pass, without free-form generation. A joint schema head scores all questions together, and the API is compatible with Jev and SystemOne. On Decision Index 0.2.1, Clef scores 98.5 on BFCL, 69.2 nDCG@10 on ToolRet, and 91.9 accuracy on API-Bank; a smaller Clef-Flash variant is also available.
- 27B model post-trained from Qwen3.8-27B with its vision encoder.
- Joint schema head scores every question option in one pass.
- Accepts text, JSON, images, and video under Apache-2.0.
- Decision Index: BFCL 98.5, ToolRet 69.2, API-Bank 91.9.
Full article1,569 words · extracted from huggingface.co · click to collapse
# Clef
- **Announcement:** [Clef decision models on the Cloudflare blog](https://blog.cloudflare.com/clef-decision-models)
- **Decision Index leaderboard:** [clef-evals.workers-ai-mle.workers.dev](https://clef-evals.workers-ai-mle.workers.dev)
Clef is a 27B multimodal model that turns a state and a schema of typed
questions into decisions. It reads the state as text, JSON, images, or video, and returns a
probability for every allowed option of every question in a single forward pass. There is no
free-form text generation and no output parsing.
The Clef API is fully compatible with Jev and SystemOne.
Clef is post-trained from [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). See
[Clef-Flash](https://huggingface.co/Cloudflare/clef-flash) for the
smaller, faster variant.
## Model
- **Backbone:** Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
- **Joint schema head:** a small transformer head that reads the backbone's final hidden states,
routes evidence from the state to each question, and scores all options of all questions jointly.
- **Output:** one logit per allowed option for each question. Apply a softmax per question to get
probabilities.
## Files
| File | Purpose |
|---|---|
| `model-*.safetensors`, `model.safetensors.index.json`, `config.json`, `generation_config.json` | Backbone, including the vision encoder |
| `joint_head.safetensors`, `joint_head_config.json` | Joint schema head |
| `joint_schema_model.py` | Record encoding, batching, the model, `load_release_model`, and `systemone` |
| `tokenizer.json`, `tokenizer_config.json`, `chat_template.jinja`, `processor_config.json` | Tokenizer and image/video processor |
| `LICENSE` | Apache-2.0 license |
## Usage
Tested with `torch` 2.11 and `transformers` 5.10.2 on a single H200. Image and video inputs also
need `pillow`.
```python
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},