Parameters
27.4B
Architecture
Hybrid attention (~75% linear / ~25% full attention), SwiGLU MLP, RoPE, RMSNorm
Released
17.09.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
—
About
Ternary-Bonsai-2-27B-gguf is the GGUF artifact of Bonsai 2 27B, Prism ML's ternary-weight compression of the Qwen3.8-27B hybrid-attention reasoning model, released September 17, 2026 under the Apache 2.0 license. It stores the full 27B-class language model (27.36B parameters total: 24.35B language backbone across 64 blocks, 2.54B embedding/LM head, 0.46B vision tower) in ternary weights at a true 1.72 bits per weight, bringing the resident model to 5.95 GB — about 9.3x smaller than the ~54 GB FP16 reference. Prism ML reports it retains 98.2% of the FP16 model's benchmark intelligence (84.78 average across 14 thinking-mode benchmarks vs 86.32 for FP16), keeping full thinking, reasoning, and agentic behavior deep in the sub-4-bit regime where conventional quantizations such as IQ2_XXS collapse toward 72.59.
Architecture
The base model architecture is unchanged from Qwen3.8-27B: a 64-block hybrid-attention causal transformer with ~75% linear attention (gated linear attention / Gated DeltaNet) and ~25% full attention interleaved in a roughly 3:1 pattern, SwiGLU MLPs, RoPE, and RMSNorm. Native context is 262,144 tokens (262K), kept practical on-device by the predominantly linear-attention backbone. The vision tower (27 blocks) ships separately as an optional Q8_0 mmproj pack (~0.63 GB) loaded only for image input, so the deployed record is multimodal-input-capable (text + image, text output).
Ternary weight format
Weights use ternary g128 quantization: every weight takes a value from {−1, 0, +1} with one shared FP16 scale per group of 128 weights, giving ~1.71 bits/weight (1.72 across the model counting the 26.2M parameters — 0.0976% of the language model, the recurrent state path of the linear-attention layers plus normalization weights — held in higher precision). Ternary coverage is end-to-end across embeddings, attention projections, MLP projections, and the LM head, with no high-precision escape hatches. Weights are stored in a rotated basis: a blockwise Hadamard rotation (block 1024, fixed ±1 signs) is folded into the stored weights offline and applied to activations at runtime; packed files declare the rotation as metadata and runtimes must apply it or refuse to load.
Two GGUF packings ship with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal, CPU): PTQ1_0 packs trits densely (1.75 bits/weight, 5.95 GB) and PQ2_0 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly and never expanded back to FP16. PQ2_0 decodes faster on H100/A100/Blackwell and processes prompts faster everywhere, PTQ1_0 decodes faster on Ada-generation cards and the L4 where memory is tightest. Stock llama.cpp cannot run these files (unknown types / silent garbage) — the PrismML-Eng llama.cpp fork is required.
Performance
Evaluated with EvalScope + vLLM on an H100 in thinking mode over 14 benchmarks, Bonsai 2 27B reaches 84.78 vs 86.32 FP16: math 96.57 (GSM8K, MATH-500, AIME25, AIME26), coding 89.42 (HumanEval+, MBPP+, LiveCodeBench), instruction following 82.66, agentic tool calling 74.92 (BFCL v3), knowledge/reasoning 79.86, vision 66.19. It outscores the conventional 2-bit IQ2_XXS build by twelve points at ~82% of that build's footprint and sits within 0.4 points of the UD-Q4_K_XL 4-bit build at one third of its footprint. Measured cross-platform throughput (batch 1): 129.9 tok/s decode on an RTX 5090 (PQ2_0), 113.9 on H100, ~28-47 tok/s on Apple M5 Pro/M5 Max laptops at 27.5 W GPU-rail draw, where the FP16 baseline does not fit at all. Its "intelligence density" (benchmark gain per GB) is 0.457/GB — ~1.8x the densest conventional build and nearly 9x FP16.
Usage
It is a reasoning model that thinks by default at xhigh reasoning effort (use medium for shorter responses; low is unsupported and behaves like xhigh). Recommended sampling: thinking mode temperature 1.0, top_p 0.95, top_k 20, min_p 0.05; instruct mode temperature 0.7, top_p 0.80, top_k 20, min_p 0.0, presence_penalty 1.5. Target use cases are laptop-local 27B agents with 262K context, privacy-sensitive/offline on-device execution, and single-GPU serving of 27B-class quality. A companion MLX build exists for native Apple Silicon inference.
Ecosystem
Maintained by Prism ML (prism-ml on Hugging Face; website prismml.com, contact@prismml.com). The Bonsai 2 line derives from Qwen/Qwen3.8-27B via quantization; previous release: Ternary Bonsai 27B (80.98 avg, 0.416 density). Community artifacts include 2 adapters, 5 finetunes, and 31 further quantizations; documentation runs through the whitepaper (bonsai-2-27b-whitepaper.pdf), the Bonsai-demo repo as the running source of truth, the MLX and mlx-swift forks, and the PrismML Discord.
Training Data Quantization of Qwen/Qwen3.8-27B; architecture unchanged from the hybrid-attention base model.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
MMLU-Redux
knowledge
|
72.65%
|
17.09.2026 |
|
GSM8K
math
|
100.00%
|
17.09.2026 |
|
HumanEval+
code
|
100.00%
|
17.09.2026 |
|
BFCL-V3
general_agent
|
100.00%
|
17.09.2026 |
|
IFBench (prompt loose)
instruction_following
|
58.97%
|
17.09.2026 |
|
OCR Bench v2
document_understanding
|
100.00%
|
17.09.2026 |
|
AIME 26
stem_reasoning
|
95.70%
|
17.09.2026 |
|
LiveCodeBench v6
stem_reasoning
|
95.32%
|
17.09.2026 |
|
MATH-500
math
|
98.53%
|
17.09.2026 |
|
MMMU-Pro
vision_language
|
79.82%
|
17.09.2026 |
|
MuSR
reasoning
|
70.63
|
17.09.2026 |
|
MBPP+
code
|
100.00%
|
17.09.2026 |
|
IFEval
instruction_following
|
93.87%
|
17.09.2026 |
|
AIME 2025
stem_reasoning
|
97.18%
|
17.09.2026 |
Architecture
- Attention
- Grouped Query Attention (24:4)
- MoE
- yes
- Layers
- 64
- Hidden size
- 5120
- Context
- 262K tokens
- RoPE θ
- 10M
- Parameters
- 27360M
Source: extracted model record · model repo
Training Pipeline
-
1
other
Ternary g128 post-training quantization
The card documents no original training or post-training of this artifact; it is a post-training quantization (PTQ) derived from Qwen/Qwen3.8-27B (architecture unchanged). Language-model weights — embeddings, attention projections, MLP projections, and LM head — are converted to ternary g128 values from {-1, 0, +1} with one shared FP16 scale per group of 128 weights, after a blockwise Hadamard rotation (block 1024, fixed ±1 signs) folded into the stored weights offline (~1.72 bits/weight overall; 26.2M parameters, 0.0976% of the language model, stay in higher precision). Two GGUF packings with custom ternary hybrid-attention kernels are shipped for llama.cpp: PTQ1_0 dense trits (1.75 bits/weight, 5.95 GB) and PQ2_0 2-bit slots (2.13 bits/weight, 7.21 GB), with the vision tower shipped separately as a Q8_0 mmproj pack. No data, learning rates, or RLHF/DPO stages are described for this artifact — the base model's training stages belong to its own record.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| MMLU-Redux | evaluation | — | text | — |
| MuSR | evaluation | — | text | — |
| GSM8K | evaluation | — | text | — |
| MATH-500 | evaluation | — | text | — |
| AIME25 | evaluation | — | text | — |
| AIME26 | evaluation | — | text | — |
| HumanEval+ | evaluation | — | text | — |
| MBPP+ | evaluation | — | text | — |
| LiveCodeBench | evaluation | — | text | — |
| IFEval | evaluation | — | text | — |
| IFBench (prompt-loose) | evaluation | — | text | — |
| BFCL v3 | evaluation | — | text | — |
| MMMU-Pro | evaluation | — | image, text | — |
| OCR Bench v2 | evaluation | — | image, text | — |
Linked Resources
Prism ML Website
https://prismml.com/
Bonsai 2 27B Whitepaper
https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdf
PrismML-Eng/Bonsai-demo
https://github.com/PrismML-Eng/Bonsai-demo
PrismML-Eng/llama.cpp
https://github.com/PrismML-Eng/llama.cpp
PrismML-Eng/mlx
https://github.com/PrismML-Eng/mlx
PrismML-Eng/mlx-swift
https://github.com/PrismML-Eng/mlx-swift
Bonsai-2 HuggingFace collection
https://huggingface.co/collections/prism-ml/bonsai-2
Companion MLX model: Ternary-Bonsai-2-27B-mlx-2bit
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit
Citation (BibTeX)
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf