Pinned PrismML @PrismML 4h Today, we’re announcing Bonsai 27B: the first 27B-class model to run on a phone. Bonsai 27...
PrismML today announced Bonsai 27B, a claimed 27B-parameter multimodal model (based on Qwen 3.6/3.7 variants) that the company says can run locally on a phone using extreme weight quantization (ternary and 1-bit variants). PrismML and community testers report ~90–95% performance retention versus the full model and fast on-device execution. This is an early signal for accelerating on-device inference but remains subject to replication, thermal/device constraints, and commercialization timelines.
Linked assets
Potential beneficiaries if on-device large-model inference becomes mainstream include mobile SoC suppliers and designers (QCOM, ARM), wafer and foundry leaders (TSM), smartphone OEMs (AAPL), and memory suppliers (MU). Cloud/API providers like MSFT face only marginal downside—on-device inference complements rather than replaces broader cloud-based AI services in the near term.
Direct leverage to AI-capable premium Android SoCs; sentiment and OEM design wins can re-rate.
IP layer exposure to broad edge-device compute growth; benefits if on-device AI becomes baseline requirement.
Its products are used in high performance computing, smartphones, Internet of things, automotive, and digital consumer electronics.
Leading-edge mobile + AI silicon content supports wafer demand; indirect but persistent tailwind.
Apple Inc.
Ecosystem/value capture if consumer expectation shifts to capable offline multimodal assistants; supports upgrade narrative.
Micron Technology, Inc.
Potential for higher DRAM/NAND content in premium phones if local models drive larger memory footprints.
Microsoft Corporation develops and supports software, services, devices, and solutions worldwide.
Only a marginal negative if local inference reduces API calls; core growth drivers remain broader than this single announcement.
Source proof
Source proof: Strong source proof | 3 extracted claims | 6 directional assets | 1 supporting author | headline-like title review
Primary sources are PrismML’s announcement thread claiming Bonsai 27B and several rapid community replies and demos reporting compressed/quantized variants (ternary, 1-bit) that run on phones with limited quality loss. Evidence is currently anecdotal and open to technical validation and reproducibility checks.
Social post thanking NVIDIA for providing DGX Spark systems to PrismML/OpenClaw engineering; hints at upcoming joint product announcements to improve local model experience and enable more local AI use-cases.
Post claims a 1-bit quantized 1.7B-parameter model (“Bonsai-1.7B”) runs at ~32 tokens/s on a OnePlus 13 using CPU-only (no GPU), implying meaningful on-device AI capability via extreme quantization.
Very low-information social post referencing “Bonsai 27b” (likely an AI model) with no concrete news, metrics, company names, or catalysts. Not actionable for public-market trading.
PrismML is promoting a demo of its “Ternary Bonsai 27B” model on Hugging Face, powered by Together Compute. This is a product/visibility update in the open-source/hosted AI inference ecosystem, but it contains no financial metrics, customer adoption data, or explicit commercial implications for public companies.
Social post praising an AI model/agent setup (“Bonsai 27b” / “bonsai 3.9gb model”) that performs tool-calls reliably in a local Hermes agent workflow. No explicit company, product vendor, revenue impact, or catalyst mentioned.
Post highlights strong performance of a local AI model (“Bonsai” by PrismML) running agentically on a consumer GPU (RTX 4070 Super), framing 2026 as a strong year for “Local AI.” It also suggests “intelligence density” as a coming benchmark driven by memory shortages, implying continued demand for efficient models, GPUs, and memory bandwidth/capacity.
Non-informational social reply expressing appreciation; contains no market, product, financial, or catalyst details.
Post highlights successful on-device LLM workflow: GLM-5.2 running locally via llama.cpp + Metal on a Mac Studio, processing a full 3-year patient chart (292 encounters) within 7.2GB using ternary quantization; data never leaves device; model constrained to asking questions. This supports a privacy-preserving, edge-compute narrative for healthcare/regulated AI workloads.
Supporting authors
The signal comprises PrismML’s official post plus corroborating community posts (testers, demos) and endorsements from technologists noting the systems-engineering achievement. No formal benchmarking, release timeline, or commercial partnership disclosures have been provided yet.
Unlock full thesis monitoring
Monitor demo/code releases and reproducibility reports, OEM or SoC vendor responses, and any performance/thermal benchmarks. Track incremental signals from QCOM, ARM, TSM, AAPL, MU, and MSFT for device design wins, memory content, and shifts in cloud-inference usage.