Bonsai 2 27B — GGUF

Prism ML

Parameters

27.4B

Architecture

Hybrid attention (~75% linear / ~25% full attention), SwiGLU MLP, RoPE, RMSNorm

Released

17.09.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal Ternary g128 (PTQ1_0, PQ2_0) Bonsai

Input Modalities

text image

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

—

Openness Index Score 100.0/100

About

Ternary-Bonsai-2-27B-gguf is the GGUF artifact of Bonsai 2 27B, Prism ML's ternary-weight compression of the Qwen3.8-27B hybrid-attention reasoning model, released September 17, 2026 under the Apache 2.0 license. It stores the full 27B-class language model (27.36B parameters total: 24.35B language backbone across 64 blocks, 2.54B embedding/LM head, 0.46B vision tower) in ternary weights at a true 1.72 bits per weight, bringing the resident model to 5.95 GB — about 9.3x smaller than the ~54 GB FP16 reference. Prism ML reports it retains 98.2% of the FP16 model's benchmark intelligence (84.78 average across 14 thinking-mode benchmarks vs 86.32 for FP16), keeping full thinking, reasoning, and agentic behavior deep in the sub-4-bit regime where conventional quantizations such as IQ2_XXS collapse toward 72.59.

Architecture

The base model architecture is unchanged from Qwen3.8-27B: a 64-block hybrid-attention causal transformer with ~75% linear attention (gated linear attention / Gated DeltaNet) and ~25% full attention interleaved in a roughly 3:1 pattern, SwiGLU MLPs, RoPE, and RMSNorm. Native context is 262,144 tokens (262K), kept practical on-device by the predominantly linear-attention backbone. The vision tower (27 blocks) ships separately as an optional Q8_0 mmproj pack (~0.63 GB) loaded only for image input, so the deployed record is multimodal-input-capable (text + image, text output).

Ternary weight format

Weights use ternary g128 quantization: every weight takes a value from {−1, 0, +1} with one shared FP16 scale per group of 128 weights, giving ~1.71 bits/weight (1.72 across the model counting the 26.2M parameters — 0.0976% of the language model, the recurrent state path of the linear-attention layers plus normalization weights — held in higher precision). Ternary coverage is end-to-end across embeddings, attention projections, MLP projections, and the LM head, with no high-precision escape hatches. Weights are stored in a rotated basis: a blockwise Hadamard rotation (block 1024, fixed ±1 signs) is folded into the stored weights offline and applied to activations at runtime; packed files declare the rotation as metadata and runtimes must apply it or refuse to load.

Two GGUF packings ship with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal, CPU): PTQ1_0 packs trits densely (1.75 bits/weight, 5.95 GB) and PQ2_0 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly and never expanded back to FP16. PQ2_0 decodes faster on H100/A100/Blackwell and processes prompts faster everywhere, PTQ1_0 decodes faster on Ada-generation cards and the L4 where memory is tightest. Stock llama.cpp cannot run these files (unknown types / silent garbage) — the PrismML-Eng llama.cpp fork is required.

Performance

Evaluated with EvalScope + vLLM on an H100 in thinking mode over 14 benchmarks, Bonsai 2 27B reaches 84.78 vs 86.32 FP16: math 96.57 (GSM8K, MATH-500, AIME25, AIME26), coding 89.42 (HumanEval+, MBPP+, LiveCodeBench), instruction following 82.66, agentic tool calling 74.92 (BFCL v3), knowledge/reasoning 79.86, vision 66.19. It outscores the conventional 2-bit IQ2_XXS build by twelve points at ~82% of that build's footprint and sits within 0.4 points of the UD-Q4_K_XL 4-bit build at one third of its footprint. Measured cross-platform throughput (batch 1): 129.9 tok/s decode on an RTX 5090 (PQ2_0), 113.9 on H100, ~28-47 tok/s on Apple M5 Pro/M5 Max laptops at 27.5 W GPU-rail draw, where the FP16 baseline does not fit at all. Its "intelligence density" (benchmark gain per GB) is 0.457/GB — ~1.8x the densest conventional build and nearly 9x FP16.

Usage

It is a reasoning model that thinks by default at xhigh reasoning effort (use medium for shorter responses; low is unsupported and behaves like xhigh). Recommended sampling: thinking mode temperature 1.0, top_p 0.95, top_k 20, min_p 0.05; instruct mode temperature 0.7, top_p 0.80, top_k 20, min_p 0.0, presence_penalty 1.5. Target use cases are laptop-local 27B agents with 262K context, privacy-sensitive/offline on-device execution, and single-GPU serving of 27B-class quality. A companion MLX build exists for native Apple Silicon inference.

Ecosystem

Maintained by Prism ML (prism-ml on Hugging Face; website prismml.com, contact@prismml.com). The Bonsai 2 line derives from Qwen/Qwen3.8-27B via quantization; previous release: Ternary Bonsai 27B (80.98 avg, 0.416 density). Community artifacts include 2 adapters, 5 finetunes, and 31 further quantizations; documentation runs through the whitepaper (bonsai-2-27b-whitepaper.pdf), the Bonsai-demo repo as the running source of truth, the MLX and mlx-swift forks, and the PrismML Discord.

Training Data Quantization of Qwen/Qwen3.8-27B; architecture unchanged from the hybrid-attention base model.

Benchmark Scores

Benchmark Score Date
MMLU-Redux
knowledge
72.65%
17.09.2026
GSM8K
math
100.00%
17.09.2026
HumanEval+
code
100.00%
17.09.2026
BFCL-V3
general_agent
100.00%
17.09.2026
IFBench (prompt loose)
instruction_following
58.97%
17.09.2026
OCR Bench v2
document_understanding
100.00%
17.09.2026
AIME 26
stem_reasoning
95.70%
17.09.2026
LiveCodeBench v6
stem_reasoning
95.32%
17.09.2026
MATH-500
math
98.53%
17.09.2026
MMMU-Pro
vision_language
79.82%
17.09.2026
MuSR
reasoning
70.63
17.09.2026
MBPP+
code
100.00%
17.09.2026
IFEval
instruction_following
93.87%
17.09.2026
AIME 2025
stem_reasoning
97.18%
17.09.2026

Architecture

Decoder Block input Embedding vocab 248K · d 5120 Full Attention GQA 24:4 · dₕ 256 ×64 MoE FFN dᴻ 17408 MTP Head ×1 speculative layer Final RMSNorm LM Head vocab 248K output
Attention
Grouped Query Attention (24:4)
MoE
yes
Layers
64
Hidden size
5120
Context
262K tokens
RoPE θ
10M
Parameters
27360M

Source: extracted model record · model repo

Type: hybrid linear attention transformer
Attention: hybrid: gated linear attention (~75% of layers) + gqa full attention (~25%, full_attention_interval=4)
Decoder: dense
Layers 64
Context length 262K
Attention heads 24
KV heads 4
Head dim 256
Hidden size 5120
Vocabulary 248K
FFN dim 17K
MTP layers 1
Vision Yes
RoPE θ 10M
Layer Pattern linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn, linear, linear, linear, full_attn
Moe {'enabled': False, 'experts_active': None, 'experts_total': None, 'first_k_dense': None, 'intermediate_size': None, 'layer_freq': None, 'route_top_k': None, 'scoring_func': None, 'shared_experts': None}
Rope Scaling {'mrope_interleaved': True, 'mrope_section': [11, 11, 10], 'partial_rotary_factor': 0.25, 'rope_type': 'default'}
Tied embeddings No

Training Pipeline

  1. 1
    other

    Ternary g128 post-training quantization

    The card documents no original training or post-training of this artifact; it is a post-training quantization (PTQ) derived from Qwen/Qwen3.8-27B (architecture unchanged). Language-model weights — embeddings, attention projections, MLP projections, and LM head — are converted to ternary g128 values from {-1, 0, +1} with one shared FP16 scale per group of 128 weights, after a blockwise Hadamard rotation (block 1024, fixed ±1 signs) folded into the stored weights offline (~1.72 bits/weight overall; 26.2M parameters, 0.0976% of the language model, stay in higher precision). Two GGUF packings with custom ternary hybrid-attention kernels are shipped for llama.cpp: PTQ1_0 dense trits (1.75 bits/weight, 5.95 GB) and PQ2_0 2-bit slots (2.13 bits/weight, 7.21 GB), with the vision tower shipped separately as a Q8_0 mmproj pack. No data, learning rates, or RLHF/DPO stages are described for this artifact — the base model's training stages belong to its own record.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
MMLU-Redux evaluation — text —
MuSR evaluation — text —
GSM8K evaluation — text —
MATH-500 evaluation — text —
AIME25 evaluation — text —
AIME26 evaluation — text —
HumanEval+ evaluation — text —
MBPP+ evaluation — text —
LiveCodeBench evaluation — text —
IFEval evaluation — text —
IFBench (prompt-loose) evaluation — text —
BFCL v3 evaluation — text —
MMMU-Pro evaluation — image, text —
OCR Bench v2 evaluation — image, text —

Related Models