PrismML has released Ternary Bonsai 2 27B, a compressed model based on Qwen3.8 27B that retains 98.2% of the full-precision model's aggregate benchmark performance while being more than 9x smaller. The model uses ternary weights with 1.76 effective bits per weight, occupying 5.9GB, and supports a 262K-token context window with multimodal input. It reaches up to 143 tokens/second on an RTX 5090 and is released under Apache 2.0.
No score is assigned. Sources and their independence are shown in the citation chain below.