Clef-Flash

Cloudflare

Parameters

9.0B

Architecture

multimodal hybrid linear attention transformer + joint schema head

Released

01.10.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Qwen3.5

Input Modalities

text image video

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

—

Openness Index Score 100.0/100

About

Clef-Flash is a 9B multimodal decision model from Cloudflare, released on October 1, 2026 under the Apache-2.0 license. It turns a state plus a schema of typed questions into decisions: Clef-Flash reads the state as text, JSON, images, or video and returns a probability for every allowed option of every question in a single forward pass - there is no free-form text generation and no output parsing. The Clef-Flash API is fully compatible with Jev and SystemOne, and Clef-Flash is post-trained from Qwen/Qwen3.5-9B. It is the smaller, faster variant of Clef (27B), and both are the first models trained by the Cloudflare Workers AI team.

Architecture

Clef-Flash consists of the Qwen3.5-9B backbone with its vision encoder (stored as standard sharded safetensors, 9B parameters, BF16) plus a joint schema head - a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly in a single forward pass. The output is one logit per allowed option for each question; a softmax per question gives probabilities. Qwen3.5-9B itself is a hybrid linear attention transformer: 32 layers in a 3:1 pattern of 24 gated linear attention layers (16 key / 32 value heads at head_dim 128) and 8 GQA full-attention layers (16 query heads / 4 KV heads at head_dim 256, partial rotary factor 0.25, interleaved MRoPE with theta 1e7 and mrope_section [11, 11, 10]), hidden size 4096, intermediate size 12288, vocab 248320, native context 262144 tokens.

Input format and typed questions

A record contains a state (any string or JSON value), optional images/video frames, and a questions mapping. Each typed question declares its allowed answers up front via type (noul true/false, choice named options, or score ordered options), optional instructions, and criteria. encode_record bounds the input at a default max_length of 16,384 tokens (with max_state_tokens), and text-only and multimodal records can be mixed in one batch. The systemone function mirrors the Jev/SystemOne POST /v1/systemone endpoint and returns answers with choice/confidence/probabilities, expected score/legend/probabilities, or the probability of true for noul.

Training and deployment

Clef-Flash is post-trained from Qwen3.5-9B (card tag post-train), with the joint schema head trained alongside it; the card documents no post-training data or hyperparameters. Tested with torch 2.11 and transformers 5.10.2 on a single H200 (pillow needed for images/videos). In Cloudflare's internal Decision Index 0.2.1 suite it is the fastest accurate decision model of the family: median latency 38.8 ms and p95 122.4 ms (Clef 27B: 209.3/238.6 ms) while leading several quality rows (BFCL 98.8, API-Bank 93.1, MMLU 91.8, ARC-Easy 99.5, WinoGrande 97.5, GPQA Diamond 51.0, ForecastBench Brier 10.6 - family best). On four Typesafe Evals business workflows (invoice processing, customer service, security incidents, agent trace observability) it scores 57.1/73.3, 77.0/76.0, 61.7, 61.7/69.8 - near Jev and ahead of Clef on customer service exact actions.

Training Data Post-trained from Qwen/Qwen3.5-9B (backbone Qwen/Qwen3.5-9B-Base with its vision encoder); no training data recipe documented.

Benchmark Scores

Benchmark Score Date
Habermas Machine
reasoning
100.00%
—
HoVer
knowledge
61.20
—
FinEntity
domain_finance
100.00%
—
HellaSwag
reasoning
100.00%
—
ARC-c
reasoning
100.00%
—
WinoGrande
reasoning
100.00%
—
MMLU
knowledge
100.00%
—
When2Call MCQ
agentic
65.60
—
GSM8K
math
49.61%
—
BRIGHT
agentic
39.30
—
BIG-Bench Hard
reasoning
69.59%
—
CRUXEval
code
86.10
—
VAST
vision_language
49.60
—
PhishNChips
cybersecurity
75.00
—
New Yorker
composite
66.10
—
BANKING77
domain_finance
90.90
—
ForecastBench
reasoning
10.60
—
API-Bank
agentic
100.00%
—
Humicroedit
reasoning
100.00%
—
GPQA Diamond
stem_reasoning
18.13%
—
POP909-CL
general_capabilities
1.60
—
NLI4CT
document_understanding
78.60
—
MMLU-Pro
knowledge
17.10%
—
CLINC150+OOS
classification
66.80
—
cfcolor
vision_language
65.80
—
BPoMP
multilingual
95.40
—
ContractNLI
document_understanding
100.00%
—
ANLI
reasoning
59.10
—
RouterBench
general_agent
100.00%
—
Amazon ESCI
domain_finance
57.40
—
CLadder
stem_reasoning
100.00%
—
SATA-Bench
general_agent
100.00%
—
ChessBench
reasoning
23.00
—
ToolRet
agentic
66.40
—
SGD/SGD-X
instruction_following
34.20
—
RAGTruth
safety
35.60
—
MuSR
reasoning
100.00%
—
ARC-e
reasoning
100.00%
—
Home appliance simulator
general_agent
100.00%
—
ACOS
general_capabilities
25.90
—

Architecture

Decoder Block ×32 input Embedding vocab 248K · d 4096 Linear / Recurrent Hybrid 16:4 · dₕ 256 ×24 Full Attention Hybrid 16:4 · dₕ 256 ×8 Dense FFN silu · d 12K Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (16:4)
Layers
32
Hidden size
4096
Context
262K tokens
Parameters
9000M

Source: Hugging Face config.json · Qwen3_5ForConditionalGeneration · exact layer pattern · model repo

Type: hybrid linear attention transformer (decoder-only backbone with vision encoder + joint schema head)
Attention: hybrid: gated linear attention (16 key heads x head_dim 128, 32 value heads x value head_dim 128) interleaved 3:1 with full GQA attention (16 query heads, 4 KV heads, head_dim 256, partial rotary 0.25, interleaved MRoPE); full_attention_interval=4
Decoder: dense
Layers 32
Context length 262K
Attention heads 16
KV heads 4
Head dim 256
Hidden size 4096
Vocabulary 248K
FFN dim 12K
MTP layers 0
Vision {'depth': 27, 'enabled': True, 'encoder': 'qwen3_5_vision', 'hidden_size': 1152, 'intermediate_size': 4304, 'num_heads': 16, 'num_position_embeddings': 2304, 'out_hidden_size': 4096, 'patch_size': 16, 'spatial_merge_size': 2, 'temporal_patch_size': 2}
RoPE θ 10M
Layer Pattern linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn
Moe {'enabled': False}
Rope Scaling {'mrope_interleaved': True, 'mrope_section': [11, 11, 10], 'partial_rotary_factor': 0.25}
Tied embeddings No
Joint Schema Head

Small transformer head (joint_head.safetensors, joint_head_config.json) that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly in a single forward pass; output is one logit per allowed option per question, softmax applied per question to get probabilities. Per model card.

Training Pipeline

  1. 1
    other

    Post-training from Qwen3.5-9B

    Clef-Flash is post-trained from Qwen/Qwen3.5-9B (itself finetuned from Qwen/Qwen3.5-9B-Base), keeping the backbone and its vision encoder. The card's 'post-train' tag and model tree ('Finetuned: Qwen/Qwen3.5-9B') identify this as a post-training step, without specifying the training data or method. As part of this step a small transformer joint schema head is trained alongside the backbone: it reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly, outputting one logit per allowed option. The card documents no training data, hyperparameters, duration, or cost.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Decision Index 0.2.1 suite evaluation — —
Typesafe Evals workflows evaluation — —

Related Models