The Tesla P4 8GB is a NVIDIA accelerator with 8 GB GDDR5 and about 192 GB/s peak memory bandwidth. AI Hashrate estimates local LLM decode (batch 1) at 8K context: fit counts weights + KV cache + 1 GB runtime overhead — not weights alone. At Q4, about 16 curated models fit fully on this card; at FP16, about 3 fit. Top Q4 speeds: MiniCPM5-1B ≈ 113.1 tok/s (estimated, fits); Qwen3.5-2B ≈ 61.1 tok/s (estimated, fits); DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx ≈ 53.1 tok/s (estimated, fits); Llama-3.2-3B-Instruct ≈ 38.2 tok/s (estimated, fits); Qwen3-4B-Instruct-2507 ≈ 30.5 tok/s (estimated, fits). Relative ranking is more reliable than absolute tok/s. See methodology for the bandwidth formula and measured-anchor policy. Methodology.