← AI Hashrate 中文

V100 16GB PCIe

NVIDIA · 16 GB HBM2 · 900 GB/s · 250 W · MSRP $8,000

CUDA

The V100 16GB PCIe is a NVIDIA accelerator with 16 GB HBM2 and about 900 GB/s peak memory bandwidth. AI Hashrate estimates local LLM decode (batch 1) at 8K context: fit counts weights + KV cache + 1 GB runtime overhead — not weights alone. At Q4, about 24 curated models fit fully on this card; at FP16, about 8 fit. Top Q4 speeds: MiniCPM5-1B ≈ 530.3 tok/s (estimated, fits); Qwen3.5-2B ≈ 286.4 tok/s (estimated, fits); DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx ≈ 249.0 tok/s (estimated, fits); Llama-3.2-3B-Instruct ≈ 179.0 tok/s (estimated, fits); gpt-oss-20b ≈ 159.1 tok/s (estimated, fits). Relative ranking is more reliable than absolute tok/s. See methodology for the bandwidth formula and measured-anchor policy. Methodology.

Table: decode speed estimates (Q4 / FP16). Measured rows override estimates. MSRP is list price, not live retail. Context default 8K.

ModelParams (B)Quanttok/sFits?
MiniCPM5-1B1.08Q4530.3 est.Yes
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1630.0Q4453.0 est.No
Qwen3-30B-A3B-Instruct-250730.5Q4412.0 est.No
GLM-4.7-Flash31.2Q4395.5 est.No
Qwen3.5-35B-A3B34.7Q4342.5 est.No
qwen35b-a3b-fable-sft-abliterated34.7Q4325.2 est.No
Ornith-1.0-35B35.0Q4320.2 est.No
Qwen3.6-35B-A3B36.0Q4319.3 est.No
Qwen3.5-2B2.0Q4286.4 est.Yes
DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx2.3Q4249.0 est.Yes
Llama-3.2-3B-Instruct3.2Q4179.0 est.Yes
gpt-oss-20b21.5Q4159.1 est.Yes
gemma-4-26B-A4B-it25.2Q4150.7 est.Yes
MiniCPM5-1B1.08FP16145.8 est.Yes
Qwen3-4B-Instruct-25074.0Q4143.2 est.Yes
NVIDIA-Nemotron-3-Nano-4B-BF164.0Q4143.2 est.Yes
Qwen3-4B-Thinking-25074.0Q4143.2 est.Yes
gemma-4-E2B-it5.1Q4112.3 est.Yes
Mistral-7B-Instruct-v0.37.2Q479.5 est.Yes
Qwen3.5-2B2.0FP1678.8 est.Yes
Qwen2.5-7B-Instruct7.6Q475.4 est.Yes
Llama-3.1-8B-Instruct8.0Q471.6 est.Yes
gemma-4-E4B-it8.0Q471.6 est.Yes
Qwen3-8B8.2Q469.8 est.Yes
DeepSeek-R1-0528-Qwen3-8B8.2Q469.8 est.Yes
Qwen3-Coder-Next79.7Q469.0 est.No
DeepSeek-R1-0528-Qwen3-8B-layer-mix-bpw-3.8-mlx2.3FP1668.5 est.Yes
Qwen3-Next-80B-A3B-Instruct81.3Q466.4 est.No
Qwen3.6-35B-A3B-NVFP421.6Q465.4 est.No
Qwen3.5-9B8.95Q464.0 est.Yes
Ornith-1.0-9B9.0Q463.6 est.Yes
Qwen3.5-27B27.0Q456.6 est.No
Mistral-7B-Instruct-v0.37.2FP1656.2 est.No
Qwen3.6-27B27.8Q452.1 est.No
Llama-3.2-3B-Instruct3.2FP1649.2 est.Yes
Devstral-Small-2-24B-Instruct-251224.0Q448.2 est.No
Qwen2.5-7B-Instruct7.6FP1648.1 est.No
gemma-4-12b-it11.95Q447.9 est.Yes
OLMo-2-1124-13B-Instruct13.7Q441.8 est.Yes
Llama-3.1-8B-Instruct8.0FP1641.5 est.No
gemma-4-E4B-it8.0FP1641.5 est.No
Qwen3-4B-Instruct-25074.0FP1639.4 est.Yes
NVIDIA-Nemotron-3-Nano-4B-BF164.0FP1639.4 est.Yes
Qwen3-4B-Thinking-25074.0FP1639.4 est.Yes
Qwen3-14B14.8Q438.7 est.Yes
Qwen2.5-14B-Instruct14.8Q438.7 est.Yes
DeepSeek-R1-Distill-Qwen-14B14.8Q438.7 est.Yes
Qwen3-8B8.2FP1638.6 est.No
DeepSeek-R1-0528-Qwen3-8B8.2FP1638.6 est.No
gemma-4-26B-A4B-it-AWQ-4bit26.6Q435.8 est.No
Qwen3.5-9B8.95FP1635.1 est.No
gemma-3-27b-it27.4Q432.8 est.No
gemma-4-E2B-it5.1FP1630.9 est.Yes
Ornith-1.0-9B9.0FP1629.5 est.No
Qwen3.5-122B-A10B-Heretic-v2-MLX-mixed-3.8bit19.8Q428.9 est.Yes
gemma-4-31B-it30.7Q423.6 est.No
gemma-4-31B-it-FP8-block31.3Q422.3 est.No
Qwen2.5-32B-Instruct32.8Q419.4 est.No
Qwen2.5-Coder-32B-Instruct32.8Q419.4 est.No
DeepSeek-R1-Distill-Qwen-32B32.8Q419.4 est.No
gpt-oss-20b21.5FP1616.0 est.No
Qwen3.6-35B-A3B-FP836.0Q414.8 est.No
gemma-4-12b-it11.95FP1612.9 est.No
gemma-4-26B-A4B-it25.2FP1611.1 est.No
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1630.0FP1610.3 est.No
Qwen3-30B-A3B-Instruct-250730.5FP169.8 est.No
GLM-4.7-Flash31.2FP169.3 est.No
OLMo-2-1124-13B-Instruct13.7FP168.6 est.No
Qwen3-14B14.8FP166.9 est.No
Qwen2.5-14B-Instruct14.8FP166.9 est.No
DeepSeek-R1-Distill-Qwen-14B14.8FP166.9 est.No
Qwen3.5-122B-A10B-Heretic-v2-MLX-mixed-3.8bit19.8FP162.9 est.No
Qwen3.6-35B-A3B-NVFP421.6FP162.3 est.No
Llama-3.1-70B-Instruct70.6Q42.1 est.No
Llama-3.3-70B-Instruct70.6Q42.1 est.No
Qwen2.5-72B-Instruct72.7Q41.9 est.No
Devstral-Small-2-24B-Instruct-251224.0FP161.7 est.No
Qwen3.5-27B27.0FP161.4 est.No
Qwen3.6-27B27.8FP161.3 est.No
gemma-4-26B-A4B-it-AWQ-4bit26.6FP161.2 est.No
gemma-3-27b-it27.4FP161.1 est.No

Fit checks

All 73 model fit checks — see the linked Fits? column in the table above