ViqusViqus
Navigate
Company
Blog
About Us
Contact
System Status
Enter Viqus Hub

PrismML Shrinks LLMs to PC/Smartphone Scale with Minimal Performance Loss

LLM compression Bonsai 2 27B Large Language Models AI reasoning Ternary weights On-device AI
September 17, 2026
Source: TechCrunch AI
Viqus Verdict Logo Viqus Verdict Logo 8
Edge AI Shift: Privacy Meets Power
Media Hype 6/10
Real Impact 8/10

Article Summary

PrismML, a startup spun out of Caltech research and advised by Ion Stoica, is making significant strides in LLM compression technology. Their latest model, Bonsai 2 27B, achieves a massive 9x-10x reduction in memory footprint compared to the original Qwen3.8 27B model, bringing it down to just 5.9 GB. This size makes it practical for deployment on consumer PCs and high-end smartphones. The core breakthrough lies in their 'ternary' weight compression technique, which drastically shrinks model weights while maintaining high benchmark parity—matching 98% of the original model's scores. Experts emphasize that running these models on-device offers substantial advantages in terms of data privacy and reduced reliance on cloud infrastructure.

Key Points

  • PrismML's new Bonsai 2 27B model is extremely small (5.9 GB) due to proprietary 'ternary' weight compression, enabling local execution on consumer hardware.
  • The compression technique maintains high performance, with the model matching 98% of the original benchmark scores, minimizing quality loss.
  • This advancement allows advanced AI models to function offline and locally on devices, significantly boosting privacy and reducing reliance on cloud APIs.

Why It Matters

The ability to run powerful, complex LLMs locally on consumer devices marks a structural shift for the industry. For businesses, this means the potential to implement advanced AI workflows without sending sensitive data to third-party cloud providers, drastically improving compliance and data privacy. It shifts the economic and technical model from API-driven cloud usage to device-centric, edge AI deployment. While other compression methods exist, PrismML's combination of high compression ratio and maintained performance puts them in a leading position to influence how consumer-grade AI applications are built and deployed.

You might also be interested in