
Startup Shrinks Reasoning AI Models to Smartphone Size Without Losing Smarts
Key takeaways
- Bonsai 2 compresses Qwen model to 5.9 GB with 98% performance retention
- Ternary weight technique reduces parameter bit depth from 16 to 3 values
- First Bonsai model downloaded 11 million times; larger models planned within months
PrismML released Bonsai 2 27B on Thursday, a compressed language model that squeezes Alibaba's Qwen down to 5.9 GB—small enough for PCs and high-end smartphones. The feat represents a 9x to 10x memory reduction while maintaining 98% of the original model's benchmark performance, up from 95% in the company's first release two months ago.
Founded by Caltech researchers and led by compression expert Babak Hassibi, PrismML uses "ternary" weights that reduce each model parameter from 16 bits to three possible values: +1, −1, or 0. This radical simplification preserves reasoning capability where competitors often lose meaningful performance. The first Bonsai model has already been downloaded 11 million times.
Hashibi plans to apply the technique to much larger models—several hundred billion parameters—within months, believing larger models will compress more effectively without intelligence loss. The startup's backing includes Khosla Ventures and Caltech, with rumors of Apple discussions still unconfirmed.
The bigger picture
PrismML enters a crowded compression field where competitors like Multiverse Computing are already well-funded, but the retained performance metric matters more than funding size. If Hassibi's prediction holds—that larger models compress even more cleanly—the implications ripple across edge AI, potentially shifting device-side inference from cloud dependency. Watch whether Apple, or other device makers, actually adopt this, and whether 100% parity becomes achievable as models scale.
We're tracking PrismML because the Caltech pedigree and Ion Stoica's involvement signal serious technical chops, not hype. Smartphone-grade LLMs aren't theoretical anymore—they're downloading by the millions. This compression story matters to anyone tired of cloud lock-in or latency.
As an Amazon Associate, LagPing earns from qualifying purchases. Product links are affiliate links.
You might also like

Inside Snorkel's Explosive Growth: How One Startup Became AI's Data Engine
20h ago

Boston Startup Summit Brings Investor Playbooks, AI Strategy, and Hiring Lessons to Founders
20h ago

Corporate-Backed Startup Factory Pivots Hard Into Physical AI With $100M War Chest
4d ago

Treble's Voice Testing Platform Lands $18M as AI Labs Rush to Perfect Audio Models
6d ago