AI Hashrate

APPLE SILICON GUIDE

Apple Silicon unified memory for local LLMs

Apple Silicon does not expose a conventional dedicated VRAM number: CPU and GPU share the memory pool. The useful input is the memory assigned to the machine and the bandwidth available to the GPU.

Unified memory changes the question

A machine can use one pool for the operating system, CPU, and GPU. Leave headroom for macOS and other applications instead of treating the printed capacity as fully available to the model.

Capacity and bandwidth are separate

A larger unified-memory configuration can hold a larger model, but bandwidth still influences decode speed. The catalog keeps both values visible so they are not confused.

Test the backend with the workload

Metal is the hardware backend in the catalog, while MLX is one of the supported runtime command paths. Recheck the fit when you change context or quantization.

Apple Silicon entries in the catalog

MachineUnified memoryBandwidthBackend
MacBook Neo A18 Pro 8GB8 GB60 GB/smetal
MacBook Pro M1 Max 32GB32 GB400 GB/smetal
Mac mini M4 32GB32 GB120 GB/smetal
MacBook Pro M1 Max 64GB64 GB400 GB/smetal
Mac mini M4 Pro 64GB64 GB273 GB/smetal
Mac Studio M3 Ultra 96GB96 GB819 GB/smetal
MacBook Pro M4 Max 128GB128 GB546 GB/smetal
Mac Studio M2 Ultra 192GB192 GB800 GB/smetal

Memory and bandwidth are catalog inputs. Actual usable memory and speed vary with system allocation, backend, thermals, and workload.

Ready to test your machine?

Enter a GPU or usable VRAM and the calculator will rank models for your workload, quantization, and context.

Open the calculator