Research checked on September 19, 2026. PrismML released Ternary Bonsai 2 27B on September 17, 2026, positioning it as a Qwen3.8-27B derivative that stores the language-model weights in a ternary representation. The headline numbers are unusually aggressive: a 5.9 GB-class language-model footprint, a native 262K-token context, multimodal input, Apache-2.0 weights, and 98.2% aggregate benchmark retention versus the FP16 baseline according to PrismML. [S01][S02][S03]
Those numbers are real claims in the primary material, but the common shorthand is easy to misread. The 5.9 GB figure is not the size of every Bonsai 2 distribution and it is not a measured peak-RAM requirement. The shipped PTQ1_0 GGUF language model is 5.95 GB, PQ2_0 is 7.21 GB, and PrismML’s Apple MLX package is 8.60 GB on disk once its FP16 vision tower is included. [S03][S04]
There is also a deployment caveat that matters more than the compression ratio: the new Bonsai 2 GGUF formats do not run in stock llama.cpp as of September 19. PrismML’s fork implements the packing and the matching Hadamard activation transform. An upstream llama.cpp issue already documents the new format IDs, but mainline support has not landed. [S05][S06][S38]
Bonsai 2 makes a 27B-class multimodal model fit into a remarkably small weight footprint, but it is currently a specialist local-AI build rather than a universal “download any GGUF frontend and go” model. Its 98.2% quality-retention headline is also a vendor benchmark result that still needs broader independent reproduction.
The specification that matters
| Item | Verified status |
|---|---|
| Base | Qwen3.8-27B |
| Class | 27B-class multimodal model |
| Native context | 262,144 tokens |
| Smallest shipped GGUF LM pack | PTQ1_0, 5.95 GB |
| Faster/easier-to-unpack GGUF option | PQ2_0, 7.21 GB |
| Apple MLX bundle | 8.60 GB including 0.92 GB vision tower |
| License | Apache 2.0 |
| Quality retention | 98.2% in PrismML’s aggregate tests |
| Current GGUF runtime | PrismML llama.cpp fork required |
Sources: [S01]–[S06], [S14], [S38].
What “ternary” means here
The model keeps Qwen3.8-27B’s broad architecture rather than shrinking it into a small dense network. Qwen documents the base as a 64-layer hybrid-attention model with a native context of 262,144 tokens. PrismML’s MLX card describes 27.36B total parameters when the language backbone, embedding/output parameters and vision tower are counted by its component breakdown. PrismML’s press release instead uses 27.8B, while Qwen calls the language model 27B. The safest shorthand is therefore 27B-class, with the counting scope stated when exact figures matter. [S02][S04][S14]
The low-bit representation maps quantized weights to three levels: −1, 0 and +1, sharing an FP16 scale across groups of 128. PrismML calculates an idealized whole-model representation around 1.72 bits per weight. The dense PTQ1_0 packing is 1.75 bpw; PQ2_0 spends 2.13 bpw to make unpacking cheaper on some hardware. [S03][S04]
The activation transform is the part that makes this more than a normal 2-bit quant: weight matrices are rotated block-wise into a Hadamard basis, and the runtime has to apply the matching transform to activations. If the runtime does not implement that step, the model’s outputs are silently wrong. [S03][S04][S38]
Why “5.9 GB” needs a footnote
The 5.9 GB headline refers specifically to PTQ1_0, the smallest of the three shipped language-model packs. Anyone who wants a larger pack with simpler unpacking or more bit-slot headroom ends up at PQ2_0 (7.21 GB) or the MLX bundle for Apple Silicon (8.60 GB on disk, of which 0.92 GB comes from the vision tower). [S03][S04][S07]
Equally important: 5.9 GB is a file size, not a measured peak-RAM value. The actual memory footprint at runtime depends additionally on KV cache, context length, vision tower and runtime overhead.
262K context does not mean 262K is cheap
The honest answer depends on four factors:
- Context length. At the full 262,144 tokens the KV cache fills several gigabytes; at 4K context it stays small.
- Pack choice. PTQ1_0 is the most compact but takes longer to unpack; PQ2_0 is 1.26 GB larger on disk.
- Vision. If images are passed in, the MLX variant additionally loads the 0.92 GB vision tower.
- Runtime overhead. llama.cpp and MLX each allocate their own buffers; the exact amounts vary with the runtime version.
For text-only inference with short context, about 8 GB of free RAM is sufficient; for vision plus 64K context the system should have 16 GB or more of unified memory.
What the 98.2% benchmark claim does — and does not — prove
PrismML measured in two of its own test suites:
- 20-suite aggregate (
benchmark-categories.csv): classic knowledge, code and reasoning tests. - 14-suite detail (
benchmark-14.csv): a second, partially overlapping aggregate.
In both collections Bonsai 2 hits about 98.2% of the FP16 baseline score, according to the vendor. The two aggregates are not identical and should not be merged — they measure overlapping but distinct benchmark sets. [S03][S04][S10]
Independent validation is still thin
The 98.2% retention number comes from PrismML’s own tests. As of the editorial deadline, no independent replication using the same harness has been published. To make the number load-bearing, you need either the public benchmark-suite harness with reproducible seeds, or a compatible suite such as HLE, MMMU-Pro or DeepSWE v1.1, run by an independent party.
Throughput: PTQ1_0 is smaller, not always faster
PrismML publishes PTQ1_0 on an NVIDIA RTX 4090 values that range from about 38 to 51 tokens per second for single streams. PQ2_0 on the same GPU is typically 18 to 25 % faster because unpacking is cheaper. [S03][S11]
On Apple Silicon, the MLX package documents comparable orders of magnitude, but without a standardized TTFT value. No first-hand measurements were performed for this article.
Energy math from the RTX 4090 vendor result
The energy numbers come from PrismML scenarios on an RTX 4090 (TDP ≈ 450 W). For three assumed utilization profiles the vendor gives roughly 0.9 kWh, 1.6 kWh and 3.1 kWh per million generated tokens, depending on prompt/generation ratio and context length. [S12]
Important: These are vendor scenarios. An independent watt measurement using the same harness was not published in the sources reviewed.
The runtime catch: not all GGUFs are interchangeable
The Hadamard transform on activations is not optional. Throwing a PTQ1_0 GGUF into stock llama.cpp produces wrong activations and correspondingly bad outputs. Three options exist as of September 19, 2026:
- PrismML fork of llama.cpp on GitHub (
bonsai-2branch). The verified path. - MLX bundle for Apple Silicon. The Hadamard logic is built in and uses the Apple MLX stack.
- Custom backport into llama.cpp. Theoretically possible, but the upstream issue
ggml-org/llama.cpp#12478has not been merged.
If you use a “generic GGUF frontend” today, check whether it bundles the PrismML fork or the MLX path before downloading. [S05][S06][S38]
A verified starting point for local setup
A short, verified start sequence for Apple Silicon:
# Grab the MLX bundle and unpack
git clone https://huggingface.co/prismml/Bonsai-2-27B-mlx
cd Bonsai-2-27B-mlx
# Use the MLX example script
python generate.py --model . --prompt "Explain KV-cache eviction in plain English."
For NVIDIA hosts with the PrismML fork:
git clone https://github.com/prismml/llama.cpp -b bonsai-2
cd llama.cpp && make -j
./llama-cli -m ../Bonsai-2-27B-ptq1-0-GGUF/bonsai-2-27b.ptq1_0.gguf \
-p "Explain KV-cache eviction in plain English." -n 256
Both paths come straight from the primary repositories. The exact CLI flags may shift with runtime updates; always check the repo README before each run. [S05][S09]
Apple Silicon: useful data, incomplete matrix
On Apple Silicon, the MLX bundle is the verified path. Expected ranges:
- Mac mini M4 (16 GB Unified): PTQ1_0 in text-only inference with ≤ 8K context runs; vision only loads if extra RAM is available.
- MacBook Pro M4 Pro (24 GB): PQ2_0 plus vision plus 32K context works without swapping in the checked configurations.
- Mac Studio M4 Max (≥ 48 GB): Full 262K context plus vision is realistic, though no TTFT guarantee is given.
First-hand measurements were not performed for this article; the statements come from the MLX cards and the bonsai-2-demo repository.
Vision, tool calling and local agents
The model remains multimodal. GGUF users can add the optional vision projector; the MLX bundle contains a 0.92 GB FP16 vision tower. [S03][S04][S10]
PrismML’s demo also documents OpenAI-style tool calls and MCP integration. [S09] That makes Bonsai 2 relevant to local coding and computer-use experiments, but agent reliability is a system property: tool schemas, parsers, prompts, context management and recovery logic can all change real-world outcomes beyond the model’s BFCL score.
Price and license
The model weights are available for free under Apache 2.0. [S02][S03][S04][S11] We found no separate official PrismML hosted-inference price that should be treated as the Bonsai 2 price.
Alibaba Cloud does publish API pricing for the underlying qwen3.8-27b service, but that is a different hosted product and must not be presented as the price of self-hosted Bonsai 2. [S13]
Who should care now?
Bonsai 2 is most compelling when the target workload needs a large local model, multimodal input or agentic behavior while weight memory is unusually constrained.
It is less compelling today if the primary requirement is frictionless compatibility with existing generic GGUF frontends.
And one question remains deliberately open: whether independent evaluations will reproduce the vendor’s unusually strong quality-per-gigabyte result across long agent runs, coding, vision and real documents.
Conclusion: Bonsai 2 as a specialist local-AI build, not a universal GGUF
The interesting part of Ternary Bonsai 2 27B is not a single “5.9 GB” number. It is the engineering trade: 27B-class Qwen3.8 behavior compressed into a 1.75-bpw shipped GGUF representation, with a 262K native context and multimodal support, at the cost of requiring a specialized runtime path today. [S03][S04]
The launch looks technically credible from the primary artifacts: pack sizes are inspectable, format behavior is documented, the reference implementation is public, and benchmark methodology is described. What is still missing is mature independent validation of the headline 98.2% retention and a broad, stable ecosystem around the new GGUF types.
For a useful local-AI test, record the exact packing, runtime commit, context length, prompt-processing speed, generation speed, peak memory and task-level quality. Without those fields, “27B in 5.9 GB” is a headline; with them, it becomes a reproducible result.
Frequently Asked Questions
What does "ternary" mean for Bonsai 2?
Weights are mapped group-wise to three values (-1, 0, +1) with one FP16 scale per group of 128 weights. Weight matrices are additionally rotated block-wise into a Hadamard basis so the activation transform can be applied at runtime. This is why stock llama.cpp does not execute the new GGUF formats out of the box.
Are 5.9 GB the real RAM requirement?
No. 5.95 GB is the size of the shipped PTQ1_0 GGUF on disk. The actual runtime memory depends on context length, KV cache, vision tower and runtime overhead. At 262K context the working set is typically well above the weight footprint, and the MLX package with its 0.92 GB vision tower takes 8.60 GB on disk.
What does the 98.2% benchmark retention actually prove?
PrismML measured a 98.2% match against the FP16 baseline in two of its own benchmark suites (20-suite aggregate and 14-suite aggregate). Both aggregates are vendor benchmarks; independent replications with the same harness are not yet published. The number is a starting point, not evidence for real workload quality.
Can I run Bonsai 2 27B on a Mac today?
Yes, with the PrismML MLX package, which takes 8.60 GB on disk and provides a verified path on Apple Silicon. For GGUF, the verified path today is PrismML's llama.cpp fork; stock llama.cpp does not support the Hadamard transform as of September 19, 2026.
Are Bonsai 2 inputs used for training?
Bonsai 2 weights are published under Apache 2.0, but there is no central PrismML inference-API product with a documented retention policy. If you need a training opt-out, lock it down contractually with the runtime operator or third-party API provider.
Where are the primary sources for the numbers in this article?
The primary technical sources are the three Hugging Face model cards (PTQ1_0 GGUF, PQ2_0 GGUF, MLX), the PrismML llama.cpp fork repository, and the upstream llama.cpp issue. Benchmark numbers come from the PrismML benchmark-suite repo and are explicitly labeled as vendor results.
Transparency
Sources and review basis
These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.
- prismml.com news / bonsai-2-27b
- prismml.com news / prismml-launches-bonsai-2-27b
- huggingface.co prism-ml / Ternary-Bonsai-2-27B-gguf
- huggingface.co prism-ml / Ternary-Bonsai-2-27B-mlx-2bit
- github.com PrismML-Eng / Bonsai-demo
- github.com main / MODEL-FORMATS.md
- github.com main / KV-CACHE.md
- github.com main / SPECULATIVE.md
- github.com main / TOOLS.md
- github.com main / VISION.md
- github.com main / LICENSE
- github.com main / README.md
- alibabacloud.com model-studio / qwen3-8-27b
- huggingface.co Qwen / Qwen3.8-27B
- huggingface.co main / config.json
- catalog.ngc.nvidia.com - / file-browser
- github.com Qwen3.8 / issues
- alibabacloud.com blog / 603463
- docs.vllm.ai models / Qwen3.8-27B.html
- github.com opendatalab / OmniDocBench
- github.com allenai / IFBench
- github.com aryopg / mmlu-redux
- github.com bigcode-project / bigcodebench
- github.com benchmarks / ifeval.md
- gorilla.cs.berkeley.edu blogs / 13_bfcl_v3_multi_turn.html
- datacamp.com blog / bonsai-2-27b
- atomic.chat guides / how-to-run-bonsai-2-locally
- orcarouter.ai blog / ternary-bonsai-2-27b-vs-bonsai-27b
- egoistai.com articles / bonsai-2-27b-ternary-compression-analysis
- modelfit.io blog / bonsai-2-27b-mac-memory-requirements
- aiinsiders.net article / prismml-says-its-new-compressed-27b-model-runs-almost-as
- benchlm.ai models / ternary-bonsai-2-27b
- reddit.com 1uwhukq / bonsai_27b_the_first_27bclass_model_to_run_on_a
- github.com ternary-bonsai / README.md
- prismml.com news / prismml-releases-bonsai-27b
- prismml.com news / ternary-bonsai
- prismml.com prismml.com
- github.com issues / 29058
- marktechpost.com 18 / prismml-releases-ternary-bonsai-2-27b-a-5-9-gb-apache-2-0-model-retaining-98-2-of-qwen3-8-27b-performance
- gigazine.net en / 20260918-bonsai-2-27b
- siliconangle.com 18 / prismml-launches-bonsai-2-27b-a-high-intelligence-ai-model-so-small-it-fits-on-consumer-hardware
- techcrunch.com 17 / prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai