activemixedx

Zain @ZainHasan6 29m This is pretty wild: • Ternary Bonsai 27B - every weight is -1, 0, or 1 and it retains 95% perf ...

PrismML’s “Bonsai 27B” posts and user tests claim a 27B-parameter multimodal model (based on Qwen3.6/3.7) can run on modern phones via aggressive weight quantization (ternary and 1-bit variants) with ~90–95% retained performance. This is an early, uncertain signal that extreme quantization and systems engineering could push more inference to edge devices, with potential implications for mobile silicon, handset features, and cloud inference mix.

Confidence
46 / 100
Assets
6
Authors
1
Outcome
open

Linked assets

Potential beneficiaries: QCOM (mobile NPUs, SoC vendors), ARM (IP licensing for mobile cores and NPUs), AAPL (on-device features and privacy-led differentiation). Potential modest negatives or mix effects for MSFT, AMZN, and NVDA if inference shifts materially off cloud, though near-term impact is likely small.

QCOMbeneficiaryopen
Confidence: 56 / 100

Mobile AI feature competition and NPU utilization are direct beneficiaries of higher on-device model capability; adoption timing uncertain.

ARMbeneficiaryopen
Confidence: 52 / 100

Broad exposure to mobile compute IP; uplift depends on whether edge inference materially increases compute intensity and silicon content.

AAPLApple Inc.beneficiaryopen

Apple Inc.

Confidence: 50 / 100

On-device AI is aligned with Apple’s privacy/latency positioning; impact depends on shipping features and consumer pull-through.

MSFTMicrosoft Corporationriskopen

Microsoft Corporation develops and supports software, services, devices, and solutions worldwide.

Confidence: 35 / 100

Any shift of inference workloads off-cloud is a narrative headwind, but enterprise AI and cloud platform momentum may offset.

AMZNAmazon.com, Inc.riskopen

Amazon.com, Inc.

Confidence: 34 / 100Start: $254.96Latest: $231.39Return: 9.24%

Similar cloud inference mix risk; magnitude likely small near-term.

NVDANVIDIA Corporationriskopen

NVIDIA Corporation operates as a data center scale AI infrastructure company.

Confidence: 33 / 100

Edge substitution is more a long-dated risk; near-term demand still driven by training and large-scale inference in datacenters.

Source proof

Source proof: Strong source proof | 5 extracted claims | 6 directional assets | 1 supporting author | headline-like title review

Primary signal: PrismML’s announcement thread claiming Bonsai 27B runs on phones; corroborating posts include user tests reporting minimal quality loss for 1-bit variants and demo comparisons versus Qwen 3.6/3.7. Other posts are qualitative praise or replies without product detail. All signals are preliminary and lack independent replication or commercial timelines.

PrismML @PrismML 4m 🦞🦞🦞 Vincent Koc @vincent_koc 3h ♥️Huge thanks to @nvidia for the DGX Sparks for @openclaw enginee...
prismml · Jul 23, 2026, 11:54 PM EDT

Social post thanking NVIDIA for providing DGX Spark systems to PrismML/OpenClaw engineering; hints at upcoming joint product announcements to improve local model experience and enable more local AI use-cases.

View source
Glenn Sonna @GlennSonna 7h Bonsai-1.7B from @PrismML: a 1-bit model decoding at 32 tok/s on a OnePlus 13. Q1_0. 237 M...
prismml · Jul 23, 2026, 3:41 PM EDT

Post claims a 1-bit quantized 1.7B-parameter model (“Bonsai-1.7B”) runs at ~32 tokens/s on a OnePlus 13 using CPU-only (no GPU), implying meaningful on-device AI capability via extreme quantization.

View source
Omead Pooladzandi @HessianFree 2h Replying to @PrismML and @togethercompute Bonsai 27b goes brrr 1 1 1 1 4 4 284 2 8 4
prismml · Jul 21, 2026, 3:10 PM EDT

Very low-information social post referencing “Bonsai 27b” (likely an AI model) with no concrete news, metrics, company names, or catalysts. Not actionable for public-market trading.

View source
PrismML @PrismML Jul 21 Try Ternary Bonsai 27B directly on Hugging Face , powered by @togethercompute huggingface.co/...
prismml · Jul 21, 2026, 3:00 PM EDT

PrismML is promoting a demo of its “Ternary Bonsai 27B” model on Hugging Face, powered by Together Compute. This is a product/visibility update in the open-source/hosted AI inference ecosystem, but it contains no financial metrics, customer adoption data, or explicit commercial implications for public companies.

View source
Omead Pooladzandi @HessianFree 13h Love to see people working with Bonsai 27b Sudo su @sudoingX 19h watch bonsai 3.9g...
prismml · Jul 19, 2026, 12:14 PM EDT

Social post praising an AI model/agent setup (“Bonsai 27b” / “bonsai 3.9gb model”) that performs tool-calls reliably in a local Hermes agent workflow. No explicit company, product vendor, revenue impact, or catalyst mentioned.

View source
Thomas Konings @tkon99 3h Finally had the time to test Bonsai by @PrismML out. On my mere RTX 4070 Super I get 45 t/s...
prismml · Jul 17, 2026, 1:41 PM EDT

Post highlights strong performance of a local AI model (“Bonsai” by PrismML) running agentically on a consumer GPU (RTX 4070 Super), framing 2026 as a strong year for “Local AI.” It also suggests “intelligence density” as a coming benchmark driven by memory shortages, implying continued demand for efficient models, GPUs, and memory bandwidth/capacity.

View source
Omead Pooladzandi @HessianFree 42m Replying to @tcarambat and @PrismML glad uve been enjoying the model
prismml · Jul 15, 2026, 9:38 AM EDT

Non-informational social reply expressing appreciation; contains no market, product, financial, or catalyst details.

View source
Maziyar PANAHI @MaziyarPanahi Jul 15 I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27...
prismml · Jul 15, 2026, 8:00 AM EDT

Post highlights successful on-device LLM workflow: GLM-5.2 running locally via llama.cpp + Metal on a Mac Studio, processing a full 3-year patient chart (292 encounters) within 7.2GB using ternary quantization; data never leaves device; model constrained to asking questions. This supports a privacy-preserving, edge-compute narrative for healthcare/regulated AI workloads.

View source

Supporting authors

Signal originates from PrismML’s pinned post and social responses (Zain @ZainHasan6, users like Aj @illetrateNerd, t.toda @Trtd6Trtd) plus commentary from technologists (Ion Stoica @istoica05). These are social-level confirmations and demonstrations rather than formal benchmarks or product releases.

Unlock full thesis monitoring

Monitor reproducibility, independent benchmarks, memory/latency/thermal metrics on target devices, and any developer SDKs or partner announcements from PrismML or handset/SoC vendors. Track handset OEM feature plans and cloud providers’ responses to evolving on-device capabilities.