New AI System Runs on Mobile Devices - ai mobile
New AI System Runs on Mobile Devices

PrismML has released Bonsai 27B, a multimodal ultra-low-bit model derived from Qwen3.6-27B, designed to enable powerful language, vision, and agentic reasoning workflows on devices where full-precision 27B models cannot practically run.

Bonsai 27B offers two compression variants: Ternary Bonsai 27B and 1-bit Bonsai 27B.

Compression Variants

Ternary Bonsai 27B uses ternary weights ({−1, 0, +1}) with FP16 group-wise scaling, yielding about 1.71 bits per weight and a deployed footprint of roughly 5.9 GB.

1-bit Bonsai 27B uses binary weights ({−1, +1}), averaging 1.125 bits per weight and occupies about 3.9 GB — compact enough to run on high-end phones.

Both variants include FP16/4-bit vision towers, which are key in supporting multimodal input such as screenshots, documents, or camera feeds, thus facilitating a wide range of applications that require the processing of visual data alongside textual information.

Multimodal Context and Reasoning

Bonsai 27B supports vision and language input, structured tool calls, agentic workflows, and sustained multi-step reasoning.

It also ships with a full context window of 262,000 tokens, enabling long document analysis, multi-turn conversation, and other prolonged‐interaction tasks.

According to the report, Bonsai 27B is designed for business owners, product teams, and professionals who need advanced AI capabilities under strict latency or privacy constraints.

For instance, on-device assistants, offline workflows, or embedded systems can significantly benefit from Bonsai 27B’s capabilities, as it allows for the deployment of 27B-class models in constrained environments without relying on cloud infrastructure.

Performance and Deployment

Across a 15-benchmark suite covering math, coding, knowledge, vision, tool-use, and instruction following, Ternary Bonsai retains around 95% of the full-precision FP16 baseline; the 1-bit variant retains about 90%.

Math and coding degrade minimally, while agentic workflows, vision, and instruction tasks show wider drops in accuracy.

Related: Crypto Platforms Drive Singapore’s Digital Asset Market

The 1-bit model enables full 27B-class reasoning on a phone, fitting within the model memory constraints of ~6 GB, including activations and KV cache.

Execution is supported via custom low-bit kernels on platforms like Apple MLX/Metal and CUDA, allowing for offline, private, or latency-sensitive settings, which is particularly useful for applications that require real-time processing or have strict data privacy requirements.

Developers and researchers can experiment with Bonsai 27B, available under the open-source Apache 2.0 license, for free download starting July 14, 2026.

PrismML provides a time-limited developer preview API for exploring its capabilities, with no subscription or usage fee stated in the official announcement.

Bonsai 27B marks a significant technical shift, pushing 27B-class model performance into the area of phones and modest laptops without wholesale reliance on the cloud.

Decision-makers should match the variant to their hardware and usage needs: use 1-bit where footprint is critical, or ternary for better overall fidelity.

Its Apache license makes experimentation and internal deployment low-risk, though the lack of published independent benchmarks in some categories means evaluations for mission-critical work should be done in-house.

For organizations balancing capability, privacy, and cost, Bonsai 27B offers a compelling local-first option, enabling them to leverage the power of 27B-class models in a more flexible and efficient manner.

Visit the official website for more.

Keep up to date with our stories on LinkedIn, Twitter, Facebook and Instagram.

Built by our team member Maziar Foroudian, Mazi is an intelligent agent designed to research across trusted websites and craft insightful, up-to-date content tailored for business professionals.