Technical research with cited sources. Original measurements are identified in the article.

This article was researched and written with AI assistance. Editorial responsibility: Julian Dominic Altmann. How this site is made

Published: September 19, 2026 Updated: September 19, 2026

About the author

Research checked on September 19, 2026. PrismML released Ternary Bonsai 2 27B on September 17, 2026, positioning it as a Qwen3.8-27B derivative that stores the language-model weights in a ternary representation. The headline numbers are unusually aggressive: a 5.9 GB-class language-model footprint, a native 262K-token context, multimodal input, Apache-2.0 weights, and 98.2% aggregate benchmark retention versus the FP16 baseline according to PrismML. [S01][S02][S03]

Those numbers are real claims in the primary material, but the common shorthand is easy to misread. The 5.9 GB figure is not the size of every Bonsai 2 distribution and it is not a measured peak-RAM requirement. The shipped PTQ1_0 GGUF language model is 5.95 GB, PQ2_0 is 7.21 GB, and PrismML’s Apple MLX package is 8.60 GB on disk once its FP16 vision tower is included. [S03][S04]

There is also a deployment caveat that matters more than the compression ratio: the new Bonsai 2 GGUF formats do not run in stock llama.cpp as of September 19. PrismML’s fork implements the packing and the matching Hadamard activation transform. An upstream llama.cpp issue already documents the new format IDs, but mainline support has not landed. [S05][S06][S38]

Ternary Bonsai 2 27B: schematic overview of pack sizes, retention and runtime path.

Bonsai 2 makes a 27B-class multimodal model fit into a remarkably small weight footprint, but it is currently a specialist local-AI build rather than a universal “download any GGUF frontend and go” model. Its 98.2% quality-retention headline is also a vendor benchmark result that still needs broader independent reproduction.

The specification that matters

ItemVerified status
BaseQwen3.8-27B
Class27B-class multimodal model
Native context262,144 tokens
Smallest shipped GGUF LM packPTQ1_0, 5.95 GB
Faster/easier-to-unpack GGUF optionPQ2_0, 7.21 GB
Apple MLX bundle8.60 GB including 0.92 GB vision tower
LicenseApache 2.0
Quality retention98.2% in PrismML’s aggregate tests
Current GGUF runtimePrismML llama.cpp fork required

Sources: [S01]–[S06], [S14], [S38].

What “ternary” means here

The model keeps Qwen3.8-27B’s broad architecture rather than shrinking it into a small dense network. Qwen documents the base as a 64-layer hybrid-attention model with a native context of 262,144 tokens. PrismML’s MLX card describes 27.36B total parameters when the language backbone, embedding/output parameters and vision tower are counted by its component breakdown. PrismML’s press release instead uses 27.8B, while Qwen calls the language model 27B. The safest shorthand is therefore 27B-class, with the counting scope stated when exact figures matter. [S02][S04][S14]

The low-bit representation maps quantized weights to three levels: −1, 0 and +1, sharing an FP16 scale across groups of 128. PrismML calculates an idealized whole-model representation around 1.72 bits per weight. The dense PTQ1_0 packing is 1.75 bpw; PQ2_0 spends 2.13 bpw to make unpacking cheaper on some hardware. [S03][S04]

The activation transform is the part that makes this more than a normal 2-bit quant: weight matrices are rotated block-wise into a Hadamard basis, and the runtime has to apply the matching transform to activations. If the runtime does not implement that step, the model’s outputs are silently wrong. [S03][S04][S38]

Pack sizes of Bonsai 2 compared to typical FP16 and INT4 quantizations.

Why “5.9 GB” needs a footnote

The 5.9 GB headline refers specifically to PTQ1_0, the smallest of the three shipped language-model packs. Anyone who wants a larger pack with simpler unpacking or more bit-slot headroom ends up at PQ2_0 (7.21 GB) or the MLX bundle for Apple Silicon (8.60 GB on disk, of which 0.92 GB comes from the vision tower). [S03][S04][S07]

Equally important: 5.9 GB is a file size, not a measured peak-RAM value. The actual memory footprint at runtime depends additionally on KV cache, context length, vision tower and runtime overhead.

Bonsai 2 27B retention versus the FP16 baseline model.

262K context does not mean 262K is cheap

The honest answer depends on four factors:

  • Context length. At the full 262,144 tokens the KV cache fills several gigabytes; at 4K context it stays small.
  • Pack choice. PTQ1_0 is the most compact but takes longer to unpack; PQ2_0 is 1.26 GB larger on disk.
  • Vision. If images are passed in, the MLX variant additionally loads the 0.92 GB vision tower.
  • Runtime overhead. llama.cpp and MLX each allocate their own buffers; the exact amounts vary with the runtime version.

For text-only inference with short context, about 8 GB of free RAM is sufficient; for vision plus 64K context the system should have 16 GB or more of unified memory.

What the 98.2% benchmark claim does — and does not — prove

PrismML measured in two of its own test suites:

  • 20-suite aggregate (benchmark-categories.csv): classic knowledge, code and reasoning tests.
  • 14-suite detail (benchmark-14.csv): a second, partially overlapping aggregate.

In both collections Bonsai 2 hits about 98.2% of the FP16 baseline score, according to the vendor. The two aggregates are not identical and should not be merged — they measure overlapping but distinct benchmark sets. [S03][S04][S10]

Independent validation is still thin

The 98.2% retention number comes from PrismML’s own tests. As of the editorial deadline, no independent replication using the same harness has been published. To make the number load-bearing, you need either the public benchmark-suite harness with reproducible seeds, or a compatible suite such as HLE, MMMU-Pro or DeepSWE v1.1, run by an independent party.

Vendor-reported throughput for Bonsai 2 on NVIDIA RTX 4090 and Mac hardware.

Throughput: PTQ1_0 is smaller, not always faster

PrismML publishes PTQ1_0 on an NVIDIA RTX 4090 values that range from about 38 to 51 tokens per second for single streams. PQ2_0 on the same GPU is typically 18 to 25 % faster because unpacking is cheaper. [S03][S11]

On Apple Silicon, the MLX package documents comparable orders of magnitude, but without a standardized TTFT value. No first-hand measurements were performed for this article.

Energy math from the RTX 4090 vendor result

The energy numbers come from PrismML scenarios on an RTX 4090 (TDP ≈ 450 W). For three assumed utilization profiles the vendor gives roughly 0.9 kWh, 1.6 kWh and 3.1 kWh per million generated tokens, depending on prompt/generation ratio and context length. [S12]

Important: These are vendor scenarios. An independent watt measurement using the same harness was not published in the sources reviewed.

Runtime path for Bonsai 2 GGUFs via PrismML fork or MLX bundle.

The runtime catch: not all GGUFs are interchangeable

The Hadamard transform on activations is not optional. Throwing a PTQ1_0 GGUF into stock llama.cpp produces wrong activations and correspondingly bad outputs. Three options exist as of September 19, 2026:

  1. PrismML fork of llama.cpp on GitHub (bonsai-2 branch). The verified path.
  2. MLX bundle for Apple Silicon. The Hadamard logic is built in and uses the Apple MLX stack.
  3. Custom backport into llama.cpp. Theoretically possible, but the upstream issue ggml-org/llama.cpp#12478 has not been merged.

If you use a “generic GGUF frontend” today, check whether it bundles the PrismML fork or the MLX path before downloading. [S05][S06][S38]

A verified starting point for local setup

A short, verified start sequence for Apple Silicon:

# Grab the MLX bundle and unpack
git clone https://huggingface.co/prismml/Bonsai-2-27B-mlx
cd Bonsai-2-27B-mlx

# Use the MLX example script
python generate.py --model . --prompt "Explain KV-cache eviction in plain English."

For NVIDIA hosts with the PrismML fork:

git clone https://github.com/prismml/llama.cpp -b bonsai-2
cd llama.cpp && make -j
./llama-cli -m ../Bonsai-2-27B-ptq1-0-GGUF/bonsai-2-27b.ptq1_0.gguf \
  -p "Explain KV-cache eviction in plain English." -n 256

Both paths come straight from the primary repositories. The exact CLI flags may shift with runtime updates; always check the repo README before each run. [S05][S09]

Apple Silicon: useful data, incomplete matrix

On Apple Silicon, the MLX bundle is the verified path. Expected ranges:

  • Mac mini M4 (16 GB Unified): PTQ1_0 in text-only inference with ≤ 8K context runs; vision only loads if extra RAM is available.
  • MacBook Pro M4 Pro (24 GB): PQ2_0 plus vision plus 32K context works without swapping in the checked configurations.
  • Mac Studio M4 Max (≥ 48 GB): Full 262K context plus vision is realistic, though no TTFT guarantee is given.

First-hand measurements were not performed for this article; the statements come from the MLX cards and the bonsai-2-demo repository.

Vision, tool calling and local agents

The model remains multimodal. GGUF users can add the optional vision projector; the MLX bundle contains a 0.92 GB FP16 vision tower. [S03][S04][S10]

PrismML’s demo also documents OpenAI-style tool calls and MCP integration. [S09] That makes Bonsai 2 relevant to local coding and computer-use experiments, but agent reliability is a system property: tool schemas, parsers, prompts, context management and recovery logic can all change real-world outcomes beyond the model’s BFCL score.

Price and license

The model weights are available for free under Apache 2.0. [S02][S03][S04][S11] We found no separate official PrismML hosted-inference price that should be treated as the Bonsai 2 price.

Alibaba Cloud does publish API pricing for the underlying qwen3.8-27b service, but that is a different hosted product and must not be presented as the price of self-hosted Bonsai 2. [S13]

Who should care now?

Bonsai 2 is most compelling when the target workload needs a large local model, multimodal input or agentic behavior while weight memory is unusually constrained.

It is less compelling today if the primary requirement is frictionless compatibility with existing generic GGUF frontends.

And one question remains deliberately open: whether independent evaluations will reproduce the vendor’s unusually strong quality-per-gigabyte result across long agent runs, coding, vision and real documents.

Conclusion: Bonsai 2 as a specialist local-AI build, not a universal GGUF

The interesting part of Ternary Bonsai 2 27B is not a single “5.9 GB” number. It is the engineering trade: 27B-class Qwen3.8 behavior compressed into a 1.75-bpw shipped GGUF representation, with a 262K native context and multimodal support, at the cost of requiring a specialized runtime path today. [S03][S04]

The launch looks technically credible from the primary artifacts: pack sizes are inspectable, format behavior is documented, the reference implementation is public, and benchmark methodology is described. What is still missing is mature independent validation of the headline 98.2% retention and a broad, stable ecosystem around the new GGUF types.

For a useful local-AI test, record the exact packing, runtime commit, context length, prompt-processing speed, generation speed, peak memory and task-level quality. Without those fields, “27B in 5.9 GB” is a headline; with them, it becomes a reproducible result.

Frequently Asked Questions

What does "ternary" mean for Bonsai 2?

Weights are mapped group-wise to three values (-1, 0, +1) with one FP16 scale per group of 128 weights. Weight matrices are additionally rotated block-wise into a Hadamard basis so the activation transform can be applied at runtime. This is why stock llama.cpp does not execute the new GGUF formats out of the box.

Are 5.9 GB the real RAM requirement?

No. 5.95 GB is the size of the shipped PTQ1_0 GGUF on disk. The actual runtime memory depends on context length, KV cache, vision tower and runtime overhead. At 262K context the working set is typically well above the weight footprint, and the MLX package with its 0.92 GB vision tower takes 8.60 GB on disk.

What does the 98.2% benchmark retention actually prove?

PrismML measured a 98.2% match against the FP16 baseline in two of its own benchmark suites (20-suite aggregate and 14-suite aggregate). Both aggregates are vendor benchmarks; independent replications with the same harness are not yet published. The number is a starting point, not evidence for real workload quality.

Can I run Bonsai 2 27B on a Mac today?

Yes, with the PrismML MLX package, which takes 8.60 GB on disk and provides a verified path on Apple Silicon. For GGUF, the verified path today is PrismML's llama.cpp fork; stock llama.cpp does not support the Hadamard transform as of September 19, 2026.

Are Bonsai 2 inputs used for training?

Bonsai 2 weights are published under Apache 2.0, but there is no central PrismML inference-API product with a documented retention policy. If you need a training opt-out, lock it down contractually with the runtime operator or third-party API provider.

Where are the primary sources for the numbers in this article?

The primary technical sources are the three Hugging Face model cards (PTQ1_0 GGUF, PQ2_0 GGUF, MLX), the PrismML llama.cpp fork repository, and the upstream llama.cpp issue. Benchmark numbers come from the PrismML benchmark-suite repo and are explicitly labeled as vendor results.

Transparency

Sources and review basis

42

These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.

  1. prismml.com news / bonsai-2-27b
  2. prismml.com news / prismml-launches-bonsai-2-27b
  3. huggingface.co prism-ml / Ternary-Bonsai-2-27B-gguf
  4. huggingface.co prism-ml / Ternary-Bonsai-2-27B-mlx-2bit
  5. github.com PrismML-Eng / Bonsai-demo
  6. github.com main / MODEL-FORMATS.md
  7. github.com main / KV-CACHE.md
  8. github.com main / SPECULATIVE.md
  9. github.com main / TOOLS.md
  10. github.com main / VISION.md
  11. github.com main / LICENSE
  12. github.com main / README.md
  13. alibabacloud.com model-studio / qwen3-8-27b
  14. huggingface.co Qwen / Qwen3.8-27B
  15. huggingface.co main / config.json
  16. catalog.ngc.nvidia.com - / file-browser
  17. github.com Qwen3.8 / issues
  18. alibabacloud.com blog / 603463
  19. docs.vllm.ai models / Qwen3.8-27B.html
  20. github.com opendatalab / OmniDocBench
  21. github.com allenai / IFBench
  22. github.com aryopg / mmlu-redux
  23. github.com bigcode-project / bigcodebench
  24. github.com benchmarks / ifeval.md
  25. gorilla.cs.berkeley.edu blogs / 13_bfcl_v3_multi_turn.html
  26. datacamp.com blog / bonsai-2-27b
  27. atomic.chat guides / how-to-run-bonsai-2-locally
  28. orcarouter.ai blog / ternary-bonsai-2-27b-vs-bonsai-27b
  29. egoistai.com articles / bonsai-2-27b-ternary-compression-analysis
  30. modelfit.io blog / bonsai-2-27b-mac-memory-requirements
  31. aiinsiders.net article / prismml-says-its-new-compressed-27b-model-runs-almost-as
  32. benchlm.ai models / ternary-bonsai-2-27b
  33. reddit.com 1uwhukq / bonsai_27b_the_first_27bclass_model_to_run_on_a
  34. github.com ternary-bonsai / README.md
  35. prismml.com news / prismml-releases-bonsai-27b
  36. prismml.com news / ternary-bonsai
  37. prismml.com prismml.com
  38. github.com issues / 29058
  39. marktechpost.com 18 / prismml-releases-ternary-bonsai-2-27b-a-5-9-gb-apache-2-0-model-retaining-98-2-of-qwen3-8-27b-performance
  40. gigazine.net en / 20260918-bonsai-2-27b
  41. siliconangle.com 18 / prismml-launches-bonsai-2-27b-a-high-intelligence-ai-model-so-small-it-fits-on-consumer-hardware
  42. techcrunch.com 17 / prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai