Ternary Bonsai 2 27B is PrismML’s “trimmed-down” large language model. It starts from Alibaba’s Qwen3.8 27B, about 27 billion parameters, which would normally take up more than 50GB at full precision; they compress the weights to only -1, 0, and +1, shrinking it to around 6–8GB while still keeping about 98% of the original’s overall benchmark performance.
In practical terms: they didn’t build an entirely new model, but rather took a 27B model that could only run on servers and compressed it so it can stay resident on a laptop too. Capabilities like math and coding drop very little; image understanding and very long multi-step tasks are a bit weaker. The trade-off is that you have to use their runtime environment, and the official Ollama still can’t run it.
It runs pretty smoothly on my M4 Pro 48GB: the weights are about 7GB, memory headroom is generous, and generation speed is roughly 20 tokens per second. It’s good enough for everyday Q&A and code edits. It’s suitable for people who don’t want to send their code out, but still want near-27B capability locally.
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit/tree/main