Thomas Konings @tkon99 3h Finally had the time to test Bonsai by @PrismML out. On my mere RTX 4070 Super I get 45 t/s...
Local AI acceleration broadens client/edge compute upgrades; memory constraints keep high-performance memory strategically valuable.
Linked assets
These are the assets attached to this thesis, along with direction, confidence, and outcome so far.
NVIDIA Corporation operates as a data center scale AI infrastructure company.
Local AI narratives can support consumer GPU demand and reinforce CUDA/software moat; still sensitive to broader AI cycle valuation.
Micron Technology, Inc.
Memory constraints theme is directly supportive; magnitude depends on HBM/DRAM pricing cycle and supply additions.
Advanced Micro Devices, Inc.
Potential beneficiary of client GPU/AI PC cycles; impact depends on competitive share vs NVDA and OEM adoption.
Edge/on-device inference aligns with product roadmap; monetization depends on AI feature pull-through and handset/PC volumes.
Microsoft Corporation develops and supports software, services, devices, and solutions worldwide.
Cloud demand could face a mix-shift if more inference moves local; counterbalanced by Copilot/software monetization.
Amazon.com, Inc.
Similar mix-shift risk to AWS inference; offset by broader cloud adoption and enterprise workloads.
Alphabet Inc.
Potential marginal inference shift to on-device; offset by Android ecosystem and model distribution advantages.
Source proof
Source proof: Strong source proof | 3 extracted claims | 7 directional assets | 1 supporting author | headline-like title review
Social post thanking NVIDIA for providing DGX Spark systems to PrismML/OpenClaw engineering; hints at upcoming joint product announcements to improve local model experience and enable more local AI use-cases.
Post claims a 1-bit quantized 1.7B-parameter model (“Bonsai-1.7B”) runs at ~32 tokens/s on a OnePlus 13 using CPU-only (no GPU), implying meaningful on-device AI capability via extreme quantization.
Very low-information social post referencing “Bonsai 27b” (likely an AI model) with no concrete news, metrics, company names, or catalysts. Not actionable for public-market trading.
PrismML is promoting a demo of its “Ternary Bonsai 27B” model on Hugging Face, powered by Together Compute. This is a product/visibility update in the open-source/hosted AI inference ecosystem, but it contains no financial metrics, customer adoption data, or explicit commercial implications for public companies.
Social post praising an AI model/agent setup (“Bonsai 27b” / “bonsai 3.9gb model”) that performs tool-calls reliably in a local Hermes agent workflow. No explicit company, product vendor, revenue impact, or catalyst mentioned.
Post highlights strong performance of a local AI model (“Bonsai” by PrismML) running agentically on a consumer GPU (RTX 4070 Super), framing 2026 as a strong year for “Local AI.” It also suggests “intelligence density” as a coming benchmark driven by memory shortages, implying continued demand for efficient models, GPUs, and memory bandwidth/capacity.
Non-informational social reply expressing appreciation; contains no market, product, financial, or catalyst details.
Post highlights successful on-device LLM workflow: GLM-5.2 running locally via llama.cpp + Metal on a Mac Studio, processing a full 3-year patient chart (292 encounters) within 7.2GB using ternary quantization; data never leaves device; model constrained to asking questions. This supports a privacy-preserving, edge-compute narrative for healthcare/regulated AI workloads.
Supporting authors
Unlock full thesis monitoring
Create an account to track this ticker thesis across linked assets, alerts, Telegram workflows, and deeper source analysis.