activemixedx

Glenn Sonna @GlennSonna 7h Bonsai-1.7B from @PrismML: a 1-bit model decoding at 32 tok/s on a OnePlus 13. Q1_0. 237 M...

Edge/on-device LLM capability is advancing via extreme quantization, benefiting mobile compute platforms more than cloud GPUs for certain consumer workloads.

Confidence
38 / 100
Assets
4
Authors
1
Outcome
open

Linked assets

These are the assets attached to this thesis, along with direction, confidence, and outcome so far.

QCOMbeneficiaryopen
Confidence: 46 / 100Start: $171.11Latest: $171.11Return: 0.00%

Mobile SoC leverage to on-device AI features; near-to-mid term catalyst sensitivity to ‘AI phone’ narratives.

ARMbeneficiaryopen
Confidence: 40 / 100Start: $283.04Latest: $283.04Return: 0.00%

CPU efficiency and ecosystem exposure if more inference runs locally; less direct than QCOM.

AAPLApple Inc.beneficiaryopen

Apple Inc.

Confidence: 34 / 100Start: $321.66Latest: $321.66Return: 0.00%

On-device AI positioning and privacy/offline benefits; impact depends on Apple’s model strategy and user adoption.

NVDANVIDIA Corporationriskopen

NVIDIA Corporation operates as a data center scale AI infrastructure company.

Confidence: 24 / 100Start: $208.76Latest: $208.76Return: 0.00%

Primarily sentiment/narrative risk for small-model inference; core datacenter training/inference demand likely less affected.

Source proof

Source proof: Strong source proof | 3 extracted claims | 4 directional assets | 1 supporting author | headline-like title review

PrismML @PrismML 4m 🦞🦞🦞 Vincent Koc @vincent_koc 3h ♥️Huge thanks to @nvidia for the DGX Sparks for @openclaw enginee...
prismml · Jul 23, 2026, 11:54 PM EDT

Social post thanking NVIDIA for providing DGX Spark systems to PrismML/OpenClaw engineering; hints at upcoming joint product announcements to improve local model experience and enable more local AI use-cases.

View source
Glenn Sonna @GlennSonna 7h Bonsai-1.7B from @PrismML: a 1-bit model decoding at 32 tok/s on a OnePlus 13. Q1_0. 237 M...
prismml · Jul 23, 2026, 3:41 PM EDT

Post claims a 1-bit quantized 1.7B-parameter model (“Bonsai-1.7B”) runs at ~32 tokens/s on a OnePlus 13 using CPU-only (no GPU), implying meaningful on-device AI capability via extreme quantization.

View source
Omead Pooladzandi @HessianFree 2h Replying to @PrismML and @togethercompute Bonsai 27b goes brrr 1 1 1 1 4 4 284 2 8 4
prismml · Jul 21, 2026, 3:10 PM EDT

Very low-information social post referencing “Bonsai 27b” (likely an AI model) with no concrete news, metrics, company names, or catalysts. Not actionable for public-market trading.

View source
PrismML @PrismML Jul 21 Try Ternary Bonsai 27B directly on Hugging Face , powered by @togethercompute huggingface.co/...
prismml · Jul 21, 2026, 3:00 PM EDT

PrismML is promoting a demo of its “Ternary Bonsai 27B” model on Hugging Face, powered by Together Compute. This is a product/visibility update in the open-source/hosted AI inference ecosystem, but it contains no financial metrics, customer adoption data, or explicit commercial implications for public companies.

View source
Omead Pooladzandi @HessianFree 13h Love to see people working with Bonsai 27b Sudo su @sudoingX 19h watch bonsai 3.9g...
prismml · Jul 19, 2026, 12:14 PM EDT

Social post praising an AI model/agent setup (“Bonsai 27b” / “bonsai 3.9gb model”) that performs tool-calls reliably in a local Hermes agent workflow. No explicit company, product vendor, revenue impact, or catalyst mentioned.

View source
Thomas Konings @tkon99 3h Finally had the time to test Bonsai by @PrismML out. On my mere RTX 4070 Super I get 45 t/s...
prismml · Jul 17, 2026, 1:41 PM EDT

Post highlights strong performance of a local AI model (“Bonsai” by PrismML) running agentically on a consumer GPU (RTX 4070 Super), framing 2026 as a strong year for “Local AI.” It also suggests “intelligence density” as a coming benchmark driven by memory shortages, implying continued demand for efficient models, GPUs, and memory bandwidth/capacity.

View source
Omead Pooladzandi @HessianFree 42m Replying to @tcarambat and @PrismML glad uve been enjoying the model
prismml · Jul 15, 2026, 9:38 AM EDT

Non-informational social reply expressing appreciation; contains no market, product, financial, or catalyst details.

View source
Maziyar PANAHI @MaziyarPanahi Jul 15 I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27...
prismml · Jul 15, 2026, 8:00 AM EDT

Post highlights successful on-device LLM workflow: GLM-5.2 running locally via llama.cpp + Metal on a Mac Studio, processing a full 3-year patient chart (292 encounters) within 7.2GB using ternary quantization; data never leaves device; model constrained to asking questions. This supports a privacy-preserving, edge-compute narrative for healthcare/regulated AI workloads.

View source

Unlock full thesis monitoring

Create an account to track this ticker thesis across linked assets, alerts, Telegram workflows, and deeper source analysis.