PrismML is building reasoning large language models designed to run locally on personal computers and smartphones. The startup has raised a $22.25 million seed round.
On Thursday, PrismML launched Bonsai 2 27B. This new model compresses Alibaba's open source Qwen3.8 27B model down to 5.9 GB. The release achieves a 9x to 10x reduction in memory compared to the original.
PrismML was founded by Caltech researchers. It is led by Babak Hassibi, a Caltech professor and compression expert. Ion Stoica serves as an adviser to the startup. Stoica co-founded Databricks and directs Berkeley's Sky Computing Lab. The company is backed by Khosla Ventures, Cerberus Capital, and Caltech.
Other companies are also working on LLM compression. Multiverse Computing is another participant in this space.
Hassibi states that PrismML achieves unique results because its models maintain nearly all original performance. Bonsai 2 matches 98 percent of Qwen's aggregate benchmark scores. The first Bonsai model, released in March, matched 95 percent of benchmark scores.
The original Bonsai model has surpassed 11 million downloads. Smaller models from PrismML have gained another 2.6 million downloads.
Hassibi notes that compression will likely always have some impact on performance. However, a 2 percent degradation is unlikely to affect actual use significantly. Software harnesses also impact overall model accuracy.
PrismML achieves this compression by shrinking model weights. Standard models use 16 bits per weight. PrismML uses a ternary weights approach that simplifies values to three options: plus one, minus one, or zero.
Stoica notes that this technology enables advanced models to run directly on user hardware. This approach provides free and private intelligence without cloud routing.
The startup's next known step is applying this compression technique to models in the several-hundred-billion-parameter range.



