“Apple Just Killed AI Subscriptions Forever” — what’s actually in it
A buy/wait/skip review of the Mac Mini M4 Pro (48 GB) as an always-on local-inference box. The headline is clickbait; the spec correction and the Bloomberg roadmap are the parts worth keeping.
The one correction the whole video hangs on
550 GB/s → 273 GB/s
Every benchmark citing 546–550 GB/s for a Mac Mini is quoting the M4 Max. The Mac Mini ships the M4 Pro — exactly half the memory bandwidth. If a reviewer quotes 550 on a Mini, they have the wrong chip or the wrong box.
Model fit on 48 GB @ 273 GB/s video’s claimed numbers
Model class
Decode
Feel
Qwen3-30B-A3B (MoE)
40–75 tok/s
fastest thing on the box
great
GPT-OSS-20B
~34 tok/s
genuinely fast
great
7–8B dense
20–30 tok/s
responsive
great
14B dense
10–20 tok/s
comfortable
fine
30–32B dense
12–18 tok/s
reads as inference, not chat
workable
70B dense
3–5 tok/s
you feel every response
wall
GPT-OSS-120B
—
needs ≥60 GB; won’t load
no fit
Escape hatches he names: 70B → M4 Max /128 GB at 12–13 tok/s. 120B-class → the AMD Strix Halo box, which he measured at 34 tok/s on GPT-OSS-120B. He concedes that round to AMD outright.
Cost of memory $ per GB of unified RAM
Box
$/GB
Read
GMKtec Evo X2 (Strix Halo, 128 GB)
$12
capacity king
Framework Desktop
$18
DGX Spark
$37
Mac Mini M4 Pro (48 GB)
$42
≈$2,399 as configured
Mac Studio M3 Ultra
$55
worst $/GB on the board
Apple loses on pure memory value and he says so. Where it wins is watts: 30–65 W under real inference load vs 45–140 W for Strix Halo and 700–800 W for an RTX tower. That’s the entire case for the Mini — a 24/7 agent loop that costs nothing to leave running.
The June 25 price move
+$200 on every M4 Pro Mac Mini config. Base went $1,399 → $1,599 overnight; the 48 GB config now ≈$2,399.
The 64 GB option was deleted entirely. Same day. Mac Studio also lost its 256 GB and 512 GB storage tiers.
Cause: Tim Cook called it a “hundred-year flood” in DRAM supply. Micron says tight beyond 2027.
AAPL fell 6% on the announcement — worst single day since April 2025.
Roadmap Mark Gurman / Bloomberg — rumor, not shipped
Oct–Nov 2026
M5 Mac MiniRumored, unconfirmed. M5 Max already ships in MacBook Pro: 128 GB, 614 GB/s. Ollama 0.19 (Mar 2026) added an MLX backend that roughly doubled decode on M5 Max.
Late 2026
Base M6 only — ~200 GB/sThe headline: M6 Pro, M6 Max and M6 Ultra are being skipped entirely. First time in the Apple Silicon era. There is no high-bandwidth M6 Mini coming.
H1 2027
M7 — fast-trackedInternal framing is reportedly “approaching Nvidia Blackwell.”
2028
M7 Ultra — 1.5 TB unified memoryPlus: M5 Ultra has reportedly been tested internally at 768 GB.
The bear case the useful half of the video
MLX’s co-creator left. Awni Hannun went to Anthropic in Feb 2026, alongside ~a dozen Apple AI researchers including the head of foundation models. The people who built the software moat are gone.
Apple pays Google ~$1B/yr for a custom Gemini to run Siri — and reportedly passed on Claude at ~$1.5B. The on-device-AI company is renting a cloud model for its flagship AI feature.
Fine-tuning on MPS is still unstable (per Sebastian Raschka). This is an inference-only box for the foreseeable future.
His verdict
Buy now
You need a 24/7 always-on inference box today.
7B–32B, or 30B-class MoE
Agentic workflows / private inference
Ollama + LM Studio just work
Go in knowing it’s 273 GB/s
Wait ~90 days
You’re not in a rush and you care about throughput.
M5 Mini rumored this fall
MLX-backend gains on M5 Max are real
Perf-per-dollar likely shifts
Skip it
70B+ is your primary workload and you need speed.
70B → M4 Max /128 GB
120B-class → Strix Halo (ROCm setup tax)
Can wait to 2027? → M7
Before you act on any of it
3:23–5:28 is a paid ad for the channel’s own “AI Master” platform — about 2 of the 17 minutes. Not content; skip it.
Zero numbers here are measured on our iron. These are his claims, on his boxes, at unstated quants and context lengths. Treat as ⚠ claimed until we run the same-prompt harness.
One number is internally broken. He says an M5 Max decode went “58 → 1112 tok/s, a 93% jump.” 58→112 is the 93%; the transcript mangled it. The 1112 figure is not real.
Everything past M5 is Gurman rumor, including the skipped-M6-Pro/Max/Ultra headline. Load-bearing if true, unconfirmed today.
Auto-caption name mangling throughout: Alama=Ollama, Quen=Qwen, Stricks Halo=Strix Halo, Rockham=ROCm, Ani Hanan=Awni Hannun.