Choose by VRAM
Model weights, KV cache, context length, and runtime overhead all consume memory. A model that fits at 4K may not fit at 32K, so the recommendation recalculates as context changes.
LOCAL LLM · HARDWARE ATLAS
Pick your hardware and workload. We put this week’s model heat against the evidence on your machine.
AI HASHRATE / WEEKLY EVIDENCE
OpenRouter heat is attention, not proof. This court checks the claim against your exact hardware, quantization, and context.
LOCAL LLM TOOL
AI Hashrate helps you choose an open-weight model from the memory you actually have. Select a GPU or enter usable VRAM to compare fit, context, estimated tokens per second, and a runnable command.
Model weights, KV cache, context length, and runtime overhead all consume memory. A model that fits at 4K may not fit at 32K, so the recommendation recalculates as context changes.
The estimate uses memory bandwidth and model size. Measured anchors are shown separately from estimates, so tok/s is a planning number rather than a benchmark promise.
After you choose a model, copy an exact command for llama.cpp, Ollama, MLX, or vLLM when the runtime is supported. Model IDs and filenames are never invented.
Fits means the estimated model weights, KV cache, and runtime overhead stay within the usable memory you entered. Leave headroom for the operating system and other applications.
They are estimates unless a card and model combination is marked measured. Actual speed varies with drivers, backend, batch size, thermals, and prompt length.
These short guides explain the constraints behind the calculator and link back to real hardware entries in the catalog.