The Titan Xp 12GB is a NVIDIA accelerator with 12 GB GDDR5X and about 548 GB/s peak memory bandwidth. AI Hashrate estimates local LLM decode (batch 1) at 8K context: fit counts weights + KV cache + 1 GB runtime overhead — not weights alone. At Q4, about 18 curated models fit fully on this card; at FP16, about 7 fit. Top Q4 speeds: MiniCPM5-1B ≈ 322.9 tok/s (estimated, fits); Qwen3.5-2B ≈ 174.4 tok/s (estimated, fits); DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx ≈ 151.6 tok/s (estimated, fits); Llama-3.2-3B-Instruct ≈ 109.0 tok/s (estimated, fits); Qwen3-4B-Instruct-2507 ≈ 87.2 tok/s (estimated, fits). Relative ranking is more reliable than absolute tok/s. See methodology for the bandwidth formula and measured-anchor policy. Methodology.