SOVEREIGN BRAIN · ARTICLE · SHEET A-07 SCALE 1:1 · REV — WIP
A·07 · ARTICLE

Bonsai 27B: A 27B-Class Brain in a 4GB Download

Bonsai 27B — PrismML's 1-bit quantized local model

The barrier for running local models just got lowered again.

Three weeks ago I wrote that Ornith-1.0-9B dropped the bar to 16GB. Now PrismML has released Bonsai 27B — and the headline number is hard to believe: a 27B-class model in a 3.9GB download. PrismML calls it the first 27B-class model to run on a phone, and the math backs them up.

What Bonsai 27B is

Bonsai 27B is the new multimodal flagship of PrismML’s Bonsai family, based on Qwen3.6 27B (the dense one). It’s aimed squarely at the workload a sovereign brain cares about: multi-step reasoning, structured tool use, long-context workflows, and coherent agentic loops.

The trick is extreme quantization. A 27B model is roughly 54GB at 16-bit precision, and even a solid 4-bit build is around 18GB — too big for a phone and for most laptops. Bonsai ships in two variants that change that completely:

VariantSizeEffective bits/weightOptimized for
Ternary Bonsai 27B5.9 GB1.71Laptop-class quality
1-bit Bonsai 27B3.9 GB1.125Phone-class footprint

Everything is open-sourced under the Apache 2.0 license.

Why I focus on the 1-bit version

The interesting one for a sovereign brain is the 1-bit build. Yes, the ternary variant keeps more quality — but the 1-bit version is smaller than Ornith’s 4-bit quant (3.9GB vs ~5.6GB) while coming from a model three times the size, and in my testing it’s super impressive for its class. Most importantly, it clears the bar that actually matters here: tool calling is good enough. It reaches for its tools and skills reliably, which — as regular readers know — is the entire ballgame for a brain that has to read and edit your files. Small enough for a phone, capable enough for the job.

That extreme compression isn’t free, and I won’t pretend a 1.125-bit model reasons like the full-precision original. But “good enough at the things a sovereign brain needs, in a 3.9GB download” is a genuinely new point on the curve. A 27B-class, multimodal, agentic-capable model now fits on hardware where, a month ago, the honest recommendation was a 9B.

What I want next

My wish for PrismML: do one based on Qwen3.6 35B A3B. I expect it would only be slightly larger than this one — but much faster, because the MoE design activates only ~3B of its 35B parameters per token. A Bonsai-style build of that model could be the best of both worlds: phone-class footprint, and speed that makes long agentic loops pleasant instead of patient.

Which one should you run?

The recommendation ladder, updated:

  • 32GB or more: run Qwen3.6 35B A3B. Still the sharpest and fastest option for a sovereign brain.
  • 16GB: run Ternary Bonsai 27B (5.9GB) for the extra quality headroom, or the 1-bit build if you want the RAM back.
  • 8GB — or a phone: run 1-bit Bonsai 27B (3.9GB). This tier simply didn’t exist before.

The floor keeps falling

This is the localmaxxing story again, on fast-forward. In June the entry ticket was a 16GB laptop. In July it’s a 3.9GB download that runs on the phone already in your pocket. The hardware didn’t change — the models keep getting more efficient, and the floor to a private, local second brain keeps quietly dropping with them.

Build your own

The tools are open source and free.

Take what you need and create your sovereign brain at sovereignbrain.me.

Project
Sovereign Brain — Homepage
Sheet
A · 01
Scale
1 : 1
Revision
Hi-Fi · Direction A