TechflierTechflierTechflier
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Bags
      • T-Shirts
    • Cart
Search
© 2025 Techflier. All Rights Reserved.
Reading: PrismML squeezes a 27B model down to laptop size
Share
Font ResizerAa
TechflierTechflier
Font ResizerAa
  • Home
  • News
  • Features
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2025 Techflier. All Rights Reserved.
News

PrismML squeezes a 27B model down to laptop size

Prism ML's Bonsai 2 uses ternary weights to shrink a 27-parameter-class model to under 6GB while keeping most of its ability.

Techflier Staff
Last updated: September 21, 2026 1:10 am
Techflier
Share
SHARE

Prism ML has launched Bonsai 2 27B, a second-generation multimodal model small enough to run on a desktop PC or a high-end phone without the usual accuracy penalty.

The trick is ternary compression, which reduces each model weight from 16 bits to three values: plus one, zero and minus one. Built on Alibaba’s Qwen3.8 27B, which occupies roughly 56GB uncompressed and needs at least 9.4GB in normal quantized form, Bonsai 2 lands at about 5.9GB while retaining around 98.2% of the original’s capabilities.

Ordinary quantization shrinks models by discarding precision, which tends to strip away knowledge and systematic ability along with the size. Prism ML’s claim is that ternary weights hold on to most of it.

Benchmarks land close to the parent model. On agentic and tool-calling work Bonsai 2 scored 77.6 against Qwen3.8’s 79.8, a gap of under three points. Coding tests across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench put it at 81.6 versus 82.2, and knowledge and reasoning measures including MMLU-Redux, GPQA Diamond and AA-LCR came in at 82.7 against 81.3.

It runs unquantized on an Nvidia GeForce GTX 5090 at 143 tokens per second, and at 46.8 tokens per second on Apple’s M5 Max chip. Prism ML says the model is roughly 40% more energy-efficient than full-precision 8B alternatives.

The practical appeal is privacy and cost: local inference keeps sensitive data off the network, and nobody pays per token. The wall is capability. Small models still lose to frontier systems on hard tasks, and narrowing that gap is what ternary compression is for.

European investors commit €300M to Open Cosmos satellite line
VC Capital Is Concentrating Faster Than Ever – What That Means for Everyone Else
AI agents take over factory purchasing as Felicis backs Magentic
Hollow-core fiber startup banks $22M to connect distant data centers
Why a $500,000 AI film at Cannes matters more than any AI demo you’ve seen
TAGGED:artificial intelligencemodel compressionon-device AIPrismMLsmall language models
SOURCES:SiliconANGLE
Share This Article
Facebook Copy Link Print
Previous Article Trustly trims a quarter of its staff in open banking reset
Next Article SEC hands tokenized stock platforms a five-year regulatory runway

Get Some Gear

 

 

 

 

Quick Links

  • News
  • Features
  • Spotlight
  • Newsletter
  • Store

About Techflier

  • About Techflier
  • Services
  • Contact Us
  • Privacy
  • Legal

Indices

TechflierTechflier
Follow US
© 2026 Techflier. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?