Apple puts a 2nm M6 in Mac mini and a 512GB-capable quad-die M5 Ultra in Mac Studio
- Apple introduced M6 in the new Mac mini and M5 Ultra in the new Mac Studio, positioning both chips for desktop performance and on-device AI workloads.
- The 2nm M6 has a 12-core CPU, 12-core GPU, Dual 16-core Neural Engine, and up to 170GB/s unified-memory bandwidth; Apple says its Neural Engine reaches up to 2x the prior generation's peak compute.
- M6's GPU places a Neural Accelerator in each of its 12 cores; Apple says peak GPU AI compute is nearly 30% higher than M5 and prompt processing for on-device LLMs is faster.
- M5 Ultra uses next-generation UltraFusion to combine four dies, the first quad-die M-series design, with up to 36 CPU cores, 80 GPU cores, and 1.2TB/s unified-memory bandwidth.
- Apple says M5 Ultra's 1.2TB/s unified-memory bandwidth is 50% above M3 Ultra, targeting large AI models and other professional desktop workloads.
Hacker News opinions
Apple was already building the hardware needed for local AI instead of joining the race to train frontier models. With 512GB of unified memory, it can run models above 100B parameters locally.
I work in AI, and Apple focusing on hardware while doing little AI software is striking. I expect local models on Macs to become normal for workflow use, and Apple may be building an LLM tuned for its own chips.
I'm looking for a cheaper, repairable Linux machine for local models. This is custom silicon rather than Nvidia, so I wonder whether Asahi support makes it practical for a company deployment.
I see 1.2TB/s as about two-thirds of an Nvidia 5090's bandwidth, but with a general-purpose computer and much more RAM. The price is still brutal.
Apple did try early with Apple Intelligence and failed badly. Letting others explore the model space instead of doubling down seems sensible.
I wish Apple would focus on local-model software, not only hardware.
There is already plenty of local inference software: Ollama, llama.cpp, LM Studio, Lemonade, and vLLM. I am not sure what Apple could add there.
Apple has MLX, and most major local-model tools support it. Apple also has RDMA to connect machines over Thunderbolt.
I disagree that Apple is only hardware. Siri AI has been under development for more than two years and is mostly local.
The 256GB memory upgrade costs £4,000 in the UK or $5,460, roughly $34 per GB. In the US it is $4,000 before sales tax, about $25 per GB, and 512GB arrives in late October.
An 80-core GPU Mac Studio with 256GB costs around $12,000, and the 512GB version may be about $14,000. That looks plausible as an on-premises inference server.
A maxed M5 Ultra Studio with 256GB and 16TB is $18,299. If 512GB costs another $6,400, it reaches about $24,699, though I still want one.
Apple is selling VRAM capacity, not ordinary RAM. Comparable 512GB HBM capacity on Nvidia hardware would cost much more, while this is a quiet desk machine with the model weights in unified memory.
I can see buyers connecting four units over RDMA for 2TB of RAM. That is cheaper than Nvidia AI hardware for some local workloads.
I do not buy the idea that many people need a 512GB Mac. The people discussing it here are unusually affluent, and high-end Apple desktops tend to have thin secondhand availability.