Back to AI
Startup Shrinks Reasoning AI Models to Smartphone Size Without Losing Smarts
AI

Startup Shrinks Reasoning AI Models to Smartphone Size Without Losing Smarts

5d ago0 views

Key takeaways

  • Bonsai 2 compresses Qwen model to 5.9 GB with 98% performance retention
  • Ternary weight technique reduces parameter bit depth from 16 to 3 values
  • First Bonsai model downloaded 11 million times; larger models planned within months

PrismML released Bonsai 2 27B on Thursday, a compressed language model that squeezes Alibaba's Qwen down to 5.9 GB—small enough for PCs and high-end smartphones. The feat represents a 9x to 10x memory reduction while maintaining 98% of the original model's benchmark performance, up from 95% in the company's first release two months ago.

Founded by Caltech researchers and led by compression expert Babak Hassibi, PrismML uses "ternary" weights that reduce each model parameter from 16 bits to three possible values: +1, −1, or 0. This radical simplification preserves reasoning capability where competitors often lose meaningful performance. The first Bonsai model has already been downloaded 11 million times.

Hashibi plans to apply the technique to much larger models—several hundred billion parameters—within months, believing larger models will compress more effectively without intelligence loss. The startup's backing includes Khosla Ventures and Caltech, with rumors of Apple discussions still unconfirmed.

The bigger picture

PrismML enters a crowded compression field where competitors like Multiverse Computing are already well-funded, but the retained performance metric matters more than funding size. If Hassibi's prediction holds—that larger models compress even more cleanly—the implications ripple across edge AI, potentially shifting device-side inference from cloud dependency. Watch whether Apple, or other device makers, actually adopt this, and whether 100% parity becomes achievable as models scale.

LagPing's take

We're tracking PrismML because the Caltech pedigree and Ion Stoica's involvement signal serious technical chops, not hype. Smartphone-grade LLMs aren't theoretical anymore—they're downloading by the millions. This compression story matters to anyone tired of cloud lock-in or latency.

Find "AI model" on Amazon

As an Amazon Associate, LagPing earns from qualifying purchases. Product links are affiliate links.

You might also like