Clef

Cloudflare

Parameters

27.0B

Architecture

Qwen3.8-27B multimodal backbone with joint schema head for typed decision outputs

Released

01.10.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Qwen3.8

Input Modalities

text image video

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

—

Openness Index Score 100.0/100

About

Clef is a 27B multimodal decision model from Cloudflare, released on October 1, 2026 under the Apache-2.0 license. It turns a state plus a schema of typed questions into decisions: Clef reads the state as text, JSON, images, or video and returns a probability for every allowed option of every question in a single forward pass - there is no free-form text generation and no output parsing. The Clef API is fully compatible with Jev and SystemOne, and Clef is post-trained from Qwen/Qwen3.8-27B. A smaller, faster variant, Clef-Flash, is also available.

Architecture

Clef consists of the Qwen3.8-27B backbone with its vision encoder (stored as standard sharded safetensors, 27B parameters, BF16) plus a joint schema head - a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly in a single forward pass. The output is one logit per allowed option for each question; a softmax per question gives probabilities. Qwen3.8-27B itself is a hybrid linear attention transformer: 64 layers in a 3:1 pattern of 48 linear-attention layers and 16 full-attention layers (gated linear attention with 16 key / 48 value heads at head_dim 128, GQA full attention with 24 query heads / 4 KV heads at head_dim 256, partial rotary factor 0.25, interleaved MRoPE with theta 1e7 and mrope_section [11, 11, 10]), hidden size 5120, intermediate size 17408, vocab 248320, native context 262144 tokens.

Input format and typed questions

A record contains a state (any string or JSON value), optional images/video frames, and a questions mapping. Each typed question declares its allowed answers up front via type (noul true/false, choice named options, or score ordered options), optional instructions, and criteria. encode_record bounds the input at a default max_length of 16,384 tokens (with max_state_tokens), and text-only and multimodal records can be mixed in one batch. The systemone function mirrors the Jev/SystemOne POST /v1/systemone endpoint and returns answers with choice/confidence/probabilities, expected score/legend/probabilities, or the probability of true for noul.

Training and deployment

Clef is post-trained from Qwen3.8-27B (card tag post-train), with the joint schema head trained alongside it; the card documents no post-training data or hyperparameters. Tested with torch 2.11 and transformers 5.10.2 on a single H200 (pillow needed for images/videos). Under the hood it leads Cloudflare's internal Decision Index 0.2.1 suite on routing/classification rows (ToolRet 69.2 nDCG@10, BANKING77 94.2 macro-F1, CLINC150+OOS 97.4, GSM8K 80.8, CRUXEval 86.7, RAGTruth 79.4) with a median latency of 209.3 ms, and leads four Typesafe Evals business workflows on invoice processing and security incidents.

Training Data Post-trained from Qwen/Qwen3.8-27B (with its vision encoder); no pretraining/post-training data recipe documented on the card.

Benchmark Scores

Benchmark Score Date
FinEntity
domain_finance
96.20
—
Habermas Machine
reasoning
68.70
—
HoVer
knowledge
100.00%
—
When2Call MCQ
agentic
100.00%
—
HellaSwag
reasoning
98.90%
—
ARC-c
reasoning
99.00%
—
BRIGHT
agentic
100.00%
—
BIG-Bench Hard
reasoning
77.84%
—
CRUXEval
code
100.00%
—
VAST
vision_language
100.00%
—
PhishNChips
cybersecurity
100.00%
—
New Yorker
composite
100.00%
—
BANKING77
domain_finance
100.00%
—
ForecastBench
reasoning
100.00%
—
GPQA Diamond
stem_reasoning
12.46%
—
POP909-CL
general_capabilities
100.00%
—
NLI4CT
document_understanding
100.00%
—
MMLU-Pro
knowledge
19.03%
—
CLINC150+OOS
classification
100.00%
—
cfcolor
vision_language
100.00%
—
BPoMP
multilingual
100.00%
—
ANLI
reasoning
100.00%
—
Amazon ESCI
domain_finance
100.00%
—
ChessBench
reasoning
100.00%
—
ToolRet
agentic
100.00%
—
SGD/SGD-X
instruction_following
100.00%
—
RAGTruth
safety
100.00%
—
ACOS
general_capabilities
100.00%
—
WinoGrande
reasoning
89.82%
—
MMLU
knowledge
95.34%
—
GSM8K
math
72.78%
—
API-Bank
agentic
91.90
—
Humicroedit
reasoning
66.70
—
ContractNLI
document_understanding
81.40
—
RouterBench
general_agent
79.70
—
CLadder
stem_reasoning
94.00
—
SATA-Bench
general_agent
33.80
—
MuSR
reasoning
83.73%
—
ARC-e
reasoning
98.11%
—
Home appliance simulator
general_agent
83.00
—

Architecture

Decoder Block ×64 input Embedding vocab 248K · d 5120 Linear / Recurrent Hybrid 24:4 · dₕ 256 ×48 Full Attention Hybrid 24:4 · dₕ 256 ×16 Dense FFN silu · d 17K Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (24:4)
Layers
64
Hidden size
5120
Context
262K tokens
Parameters
27000M

Source: Hugging Face config.json · Qwen3_5ForConditionalGeneration · exact layer pattern · model repo

Type: hybrid linear attention transformer (Qwen3.8-27B backbone + joint schema transformer head)
Attention: gated linear attention (48 layers, 16 key and 48 value heads, head_dim 128) + GQA full attention (16 layers, 24 query heads, 4 kv heads, head_dim 256)
Decoder: hybrid
Layers 64
Context length 262K
Attention heads 24
KV heads 4
Head dim 256
Hidden size 5120
Vocabulary 248K
FFN dim 17K
MTP layers 0
Vision {'deepstack_visual_indexes': [], 'depth': 27, 'enabled': True, 'encoder': 'qwen3_5_vision', 'hidden_size': 1152, 'intermediate_size': 4304, 'num_heads': 16, 'out_hidden_size': 5120, 'patch_size': 16, 'spatial_merge_size': 2, 'temporal_patch_size': 2}
RoPE θ 10M
Layer Pattern 3:1 pattern x 16 (linear, linear, linear, full_attn)
Moe {'enabled': False}
Rope Scaling {'mrope_interleaved': True, 'mrope_section': [11, 11, 10], 'partial_rotary_factor': 0.25, 'rope_type': 'default'}
Tied embeddings No
Joint Schema Head

Small transformer head stored separately (joint_head.safetensors, joint_head_config.json). It reads the backbone final hidden states, routes evidence from the state to each question in the typed schema, and scores all options of all questions jointly in a single forward pass (one logit per allowed option, softmax per question). This is task-head evidence routing, not MoE expert routing.

Training Pipeline

  1. 1
    other

    Post-training from Qwen3.8-27B

    The model card states Clef is post-trained from Qwen/Qwen3.8-27B (tagged `post-train`). The released weights include the Qwen3.8-27B backbone with its vision encoder plus a separately stored joint schema head — a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly to produce one logit per allowed option. The card does not document the post-training data, method, or any hyperparameters.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Decision Index 0.2.1 evaluation — text, image, video —
Typesafe Evals evaluation 4 end-to-end business workflows —

Related Models