PrismML hopes its tiny LLM will change how we all use AI

PrismML, a startup founded by Caltech researchers and led by professor Babak Hassibi, is producing highly compressed large language models that can run on personal devices. Its latest release, Bonsai 2 27B, reduces Alibaba’s Qwen3.8 27B to 5.9 GB while retaining about 98% of the original model’s benchmark performance.

By AI Newsroom· Reviewed by Pranav, Founder & Editor-in-ChiefPublished about 10 hours agoUpdated about 3 hours ago0 views
PrismML hopes its tiny LLM will change how we all use AI

Why It Matters

If PrismML’s compression preserves model capability at much smaller sizes, it could enable advanced LLMs to run locally on phones and PCs, changing deployment, privacy, and cost dynamics for AI services. The company’s approach may shift how developers and users access high-performance models outside cloud infrastructure.

Key Facts

  • Company: PrismML
  • Founders / leadership: Founded by Caltech researchers; CEO Babak Hassibi
  • Adviser: Ion Stoica
  • Funding: $22.25 million seed round
  • Investors: Khosla Ventures, Cerberus Capital, Caltech

PrismML is marketing a compression-first approach to large language models, arguing that high reasoning performance need not require gigantic models. The startup’s newest public model, Bonsai 2 27B, compresses Alibaba’s open-source Qwen3.8 27B down to approximately 5.9 GB, a roughly 9x–10x reduction in memory footprint versus the original. PrismML says Bonsai 2 preserves about 98% of Qwen’s aggregate benchmark scores, an improvement over its earlier Bonsai release that matched roughly 95%. The company reports that the original Bonsai has been downloaded more than 11 million times, and its even smaller models have added another 2.6 million downloads. The firm’s technical approach centers on shrinking model weights. Whereas weights are typically stored with 16-bit precision, PrismML applies a ternary scheme that restricts values to +1, −1, or 0, dramatically lowering storage needs. The company provides further technical details on its Hugging Face project page. PrismML was founded by Caltech researchers and is led by Hassibi, an expert in compression techniques; Ion Stoica of Berkeley’s Sky Computing Lab serves as an adviser. The startup has drawn investor support from Khosla Ventures, Cerberus Capital, and Caltech. Hassibi declined to comment on rumors that PrismML is in discussions with Apple. Looking ahead, PrismML aims to apply its compression method to much larger models, targeting releases in the several-hundred-billion-parameter range within months. Company leaders argue that larger models may offer more room for compression with less loss of capability, and backers note potential benefits such as on-device availability, lower ongoing cloud costs, and improved privacy when models run locally rather than in the cloud.

Keep Reading