The V100 16GB PCIe is a NVIDIA accelerator with 16 GB HBM2 and about 900 GB/s peak memory bandwidth. AI Hashrate estimates local LLM decode (batch 1) at 8K context: fit counts weights + KV cache + 1 GB runtime overhead — not weights alone. At Q4, about 24 curated models fit fully on this card; at FP16, about 8 fit. Top Q4 speeds: MiniCPM5-1B ≈ 530.3 tok/s (estimated, fits); Qwen3.5-2B ≈ 286.4 tok/s (estimated, fits); DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx ≈ 249.0 tok/s (estimated, fits); Llama-3.2-3B-Instruct ≈ 179.0 tok/s (estimated, fits); gpt-oss-20b ≈ 159.1 tok/s (estimated, fits). Relative ranking is more reliable than absolute tok/s. See methodology for the bandwidth formula and measured-anchor policy. Methodology.