Zain @ZainHasan6 29m This is pretty wild: • Ternary Bonsai 27B - every weight is -1, 0, or 1 and it retains 95% perf ...
PrismML’s “Bonsai 27B” posts and user tests claim a 27B-parameter multimodal model (based on Qwen3.6/3.7) can run on modern phones via aggressive weight quantization (ternary and 1-bit variants) with ~90–95% retained performance. This is an early, uncertain signal that extreme quantization and systems engineering could push more inference to edge devices, with potential implications for mobile silicon, handset features, and cloud inference mix.
Linked assets
Potential beneficiaries: QCOM (mobile NPUs, SoC vendors), ARM (IP licensing for mobile cores and NPUs), AAPL (on-device features and privacy-led differentiation). Potential modest negatives or mix effects for MSFT, AMZN, and NVDA if inference shifts materially off cloud, though near-term impact is likely small.
Mobile AI feature competition and NPU utilization are direct beneficiaries of higher on-device model capability; adoption timing uncertain.
Broad exposure to mobile compute IP; uplift depends on whether edge inference materially increases compute intensity and silicon content.
Apple Inc.
On-device AI is aligned with Apple’s privacy/latency positioning; impact depends on shipping features and consumer pull-through.
Microsoft Corporation develops and supports software, services, devices, and solutions worldwide.
Any shift of inference workloads off-cloud is a narrative headwind, but enterprise AI and cloud platform momentum may offset.
Amazon.com, Inc.
Similar cloud inference mix risk; magnitude likely small near-term.
NVIDIA Corporation operates as a data center scale AI infrastructure company.
Edge substitution is more a long-dated risk; near-term demand still driven by training and large-scale inference in datacenters.
Source proof
Source proof: Strong source proof | 5 extracted claims | 6 directional assets | 1 supporting author | headline-like title review
Primary signal: PrismML’s announcement thread claiming Bonsai 27B runs on phones; corroborating posts include user tests reporting minimal quality loss for 1-bit variants and demo comparisons versus Qwen 3.6/3.7. Other posts are qualitative praise or replies without product detail. All signals are preliminary and lack independent replication or commercial timelines.
Social post thanking NVIDIA for providing DGX Spark systems to PrismML/OpenClaw engineering; hints at upcoming joint product announcements to improve local model experience and enable more local AI use-cases.
Post claims a 1-bit quantized 1.7B-parameter model (“Bonsai-1.7B”) runs at ~32 tokens/s on a OnePlus 13 using CPU-only (no GPU), implying meaningful on-device AI capability via extreme quantization.
Very low-information social post referencing “Bonsai 27b” (likely an AI model) with no concrete news, metrics, company names, or catalysts. Not actionable for public-market trading.
PrismML is promoting a demo of its “Ternary Bonsai 27B” model on Hugging Face, powered by Together Compute. This is a product/visibility update in the open-source/hosted AI inference ecosystem, but it contains no financial metrics, customer adoption data, or explicit commercial implications for public companies.
Social post praising an AI model/agent setup (“Bonsai 27b” / “bonsai 3.9gb model”) that performs tool-calls reliably in a local Hermes agent workflow. No explicit company, product vendor, revenue impact, or catalyst mentioned.
Post highlights strong performance of a local AI model (“Bonsai” by PrismML) running agentically on a consumer GPU (RTX 4070 Super), framing 2026 as a strong year for “Local AI.” It also suggests “intelligence density” as a coming benchmark driven by memory shortages, implying continued demand for efficient models, GPUs, and memory bandwidth/capacity.
Non-informational social reply expressing appreciation; contains no market, product, financial, or catalyst details.
Post highlights successful on-device LLM workflow: GLM-5.2 running locally via llama.cpp + Metal on a Mac Studio, processing a full 3-year patient chart (292 encounters) within 7.2GB using ternary quantization; data never leaves device; model constrained to asking questions. This supports a privacy-preserving, edge-compute narrative for healthcare/regulated AI workloads.
Supporting authors
Signal originates from PrismML’s pinned post and social responses (Zain @ZainHasan6, users like Aj @illetrateNerd, t.toda @Trtd6Trtd) plus commentary from technologists (Ion Stoica @istoica05). These are social-level confirmations and demonstrations rather than formal benchmarks or product releases.
Unlock full thesis monitoring
Monitor reproducibility, independent benchmarks, memory/latency/thermal metrics on target devices, and any developer SDKs or partner announcements from PrismML or handset/SoC vendors. Track handset OEM feature plans and cloud providers’ responses to evolving on-device capabilities.