PrismML Shrinks LLMs to PC/Smartphone Scale with Minimal Performance Loss
8
What is the Viqus Verdict?
We evaluate each news story based on its real impact versus its media hype to offer a clear and objective perspective.
AI Analysis:
The underlying technology (on-device LLMs) is a high-impact structural shift, but the current release is an incremental step (27B model), keeping the hype score moderate.
Article Summary
PrismML, a startup spun out of Caltech research and advised by Ion Stoica, is making significant strides in LLM compression technology. Their latest model, Bonsai 2 27B, achieves a massive 9x-10x reduction in memory footprint compared to the original Qwen3.8 27B model, bringing it down to just 5.9 GB. This size makes it practical for deployment on consumer PCs and high-end smartphones. The core breakthrough lies in their 'ternary' weight compression technique, which drastically shrinks model weights while maintaining high benchmark parity—matching 98% of the original model's scores. Experts emphasize that running these models on-device offers substantial advantages in terms of data privacy and reduced reliance on cloud infrastructure.Key Points
- PrismML's new Bonsai 2 27B model is extremely small (5.9 GB) due to proprietary 'ternary' weight compression, enabling local execution on consumer hardware.
- The compression technique maintains high performance, with the model matching 98% of the original benchmark scores, minimizing quality loss.
- This advancement allows advanced AI models to function offline and locally on devices, significantly boosting privacy and reducing reliance on cloud APIs.

