Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Two months after the initial release of Bonsai 27B, PrismML introduces Ternary Bonsai 2 27B, a more capable model with a smaller footprint. This new model is based on Qwen3.8 27B and offers improved reasoning, coding, vision, and agentic capabilities while maintaining the benefits of a smaller memory footprint, high local throughput, and better energy efficiency.
Ternary Bonsai 2 27B utilizes ternary weights with FP16 group-wise scaling, resulting in 1.76 effective bits per weight and a total model size of 5.9GB. Despite being 9x smaller than the full-precision counterpart, it retains 98.2% of aggregate benchmark performance. This compression enables the model to run efficiently on local devices, making it suitable for a wider range of applications.
Compared to the previous Bonsai 27B generation, Ternary Bonsai 2 27B shows significant improvements in coding agents, tool-use systems, multimodal workflows, and long-horizon tasks. It also demonstrates higher throughput and better energy efficiency on various hardware configurations, including NVIDIA GeForce RTX 5090, M5 Max, and RTX 4090.
With its improved intelligence density and better energy efficiency, Ternary Bonsai 2 27B paves the way for local models to handle real knowledge work, including coding-agent loops, computer-use workflows, private document analysis, and multimodal debugging. This advancement in low-bit models has the potential to change the economics and architecture of AI systems across devices, workstations, and datacenters, enabling more efficient deployment and utilization of AI across various platforms.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.