Apple's M5 Ultra Mac Studio reaches 512GB unified memory for local LLMs and four-node AI clusters
- M5 Ultra Mac Studio scales to a 36-core CPU, an 80-core GPU, 512GB unified memory, and 1.2TB/s memory bandwidth, which Apple says permits enormous LLMs to run entirely on device.
- Apple says Neural Accelerators built into every GPU core raise M5 Ultra AI performance to 4.3 times that of M3 Ultra and 9.8 times that of M1 Ultra.
- Four Mac Studio systems can use Thunderbolt 5 and RDMA to create a shared memory pool, with Apple claiming up to 3 times the AI inference performance of one system.
- The M5 Max model has an 18-core CPU, up to 40 GPU cores, and up to 128GB unified memory; the M5 Ultra starts at $5,499 in the US, while M5 Max starts at $2,499.
- Apple introduced Core AI, a macOS framework for building, running, and deploying models on Apple silicon across unified memory, CPU, GPU, and Neural Engine; MLX remains its open-source machine learning framework.
Hacker News opinions
The M5 Ultra with a 64-core GPU and 96GB RAM costs €6,649 in Europe. Ouch.
US prices usually exclude VAT, while EU prices include it, so the direct comparison is misleading.
I am surprised there is no 1TB RAM configuration for people with excessive budgets or VC money.
Apple may be limiting large-memory options because a 1TB configuration could replace sales of multiple machines. Memory supply constraints may also be involved, since high-capacity M3 Ultra options disappeared and 128GB MacBook Pro orders reportedly take six weeks or longer.
The 512GB unified-memory option is due in October. If 256GB costs an extra $4,000, a 512GB setup could land near $20,000.
512GB unified memory should be excellent for local LLMs. I take unified RAM to mean the CPU and GPU share one memory pool rather than having separate system RAM and VRAM.
170GB/s is slow for local LLM work. That is the Mac mini, while the Mac Studio reaches 1.2TB/s.
Compared with VRAM it is slow, but 170GB/s is still fast by normal computer standards for an entry-level chip. Apple reserves much higher bandwidth for Pro, Max, and Ultra parts.
I am considering a Studio instead of a docked MacBook Pro. I rarely need portability, so a desktop plus a Neo for travel might make more sense.
I have run a Studio for two years and it has been great. Lack of portability matters mostly during travel, and I can just pause work then.
My maxed M3 Max MacBook Pro got hot and loud under local LLMs, and its battery degraded from heat. My M4 Max Studio runs Gemma 3/4 and gpt-oss 120b all day without audible fan noise, though my M5 Max MacBook Pro gets about 100 tokens/s on Gemma 4 27B before overheating after a few minutes.
I want an Ultra Studio as a long-lived server, so I am holding out. I keep an M1 Pro always on and use Tailscale to reach it remotely, but a Studio would be a better compute hub.
A 256GB-memory configuration is around $10,000, and 512GB may be around $20,000. It will not handle models above 1T parameters comfortably, but Thunderbolt 5 and three ports permit a fully connected four-machine cluster.
An RTX 6000 has 96GB at 1.7TB/s and costs $13,000. A Mac with 256GB at 1.2TB/s looks competitive by comparison.
I would pay for 1TB unified memory because 4-bit quantized models above 1T parameters need it reliably. 512GB is only barely enough for the models I want, and I cannot justify an M5 Ultra without the 1TB option.
I do not expect 1TB to be cheap. With 256GB already above $10,000, 1TB would likely be around $20,000.