US National WireUS NATIONAL WIRE
TechOpinion

PrismML's Compression Play: A High-Value Target for Hyperscalers

Portrait of Owen Pearce
Owen PearceM&A / IPOs / exitsSep 17AI
PrismML's Compression Play: A High-Value Target for Hyperscalers

AI-generated image · US National Wire

By shrinking model weights to ternary values, the Caltech-led startup is slashing memory requirements—a critical lever for reducing inference costs at scale.

Opinion: In the current AI arms race, the primary bottleneck for hyperscalers is no longer just training capacity, but the staggering cost of inference. As these giants seek to optimize their margins, PrismML is positioning itself as a critical architectural solution. By focusing on extreme efficiency rather than raw parameter growth, the startup is creating a blueprint for deploying high-performance reasoning models at a fraction of the traditional hardware cost.

As first reported by TechCrunch, PrismML is utilizing a "ternary" weight approach to compress large language models (LLMs). While standard weights typically require 16 bits, PrismML simplifies these to just three values: +1, -1, or 0. This mechanism allows the company to drastically reduce the memory footprint of a model. A primary example is the recently released Bonsai 2 27B, which compresses Alibaba's open-source Qwen3.8 27B model down to 5.9 GB. TechCrunch reports this represents a 9x to 10x reduction in memory compared to the original version.

From a deals perspective, the value proposition is not just the size reduction, but the retention of intelligence. CEO Babak Hassibi told TechCrunch that Bonsai 2 matches 98% of the aggregate benchmark scores of the original Qwen model. This is a notable improvement over the first Bonsai model released in March, which matched 95%. Hassibi argues that a 2% degradation is unlikely to meaningfully impact actual performance in real-world use, as uncompressed models are not perfectly accurate and benchmarks are not flawless reflections of task performance.

PrismML's roadmap suggests an even greater opportunity for acquisition by cloud providers. Hassibi informed TechCrunch that the company intends to apply this compression to models in the several-hundred-billion-parameter range in the coming months, noting that larger models provide more room to compress without losing intelligence.

While the company is still early-stage—having raised a $22.25 million seed round from investors including Cerberus Capital, Khosla Ventures, and Caltech—its pedigree is significant. The startup was established by a group of Caltech researchers and is led by professor Babak Hassibi, with Ion Stoica serving as an advisor. Stoica, a co-founder of Databricks and director of Berkeley’s Sky Computing Lab, told TechCrunch that this technology enables advanced models to run privately and for free on a user's own device.

Although TechCrunch noted that other firms, such as Multiverse Computing, are also pursuing LLM compression, PrismML's focus on maintaining near-total performance parity makes it an attractive target. While CEO Babak Hassibi declined to comment on rumors regarding talks with Apple, the ability to move reasoning models onto PCs and smartphones is a strategic imperative for any hardware or cloud giant looking to slash the overhead of AI inference.

Sources

More from Owen Pearce