Both list the same memory bandwidth (900 GB/s), so per-token decode speed should be similar whenever a model fits both. The V100 32GB SXM2 fits more of the catalog at Q4/8K — 43 vs 25 of 73 models (32 vs 16 GB VRAM). Largest model that fits the V100 32GB SXM2 but not the V100 16GB PCIe: Qwen3.6-35B-A3B-FP8 (36.0B). MSRP is $8,000 for the V100 16GB PCIe vs $10,000 for the V100 32GB SXM2. MSRP is a launch list price, not live retail — Amazon prices move; click through for the current price. Not retail-shoppable: V100 16GB PCIe, V100 32GB SXM2 (datacenter-class; cloud rental is the realistic way to use it). In the Q4/8K head-to-head below (49 shared models), the V100 32GB SXM2 posts the higher tok/s in 15 rows vs 9. Methodology.
Spec comparison
V100 16GB PCIe
V100 32GB SXM2
Vendor
NVIDIA
NVIDIA
VRAM
16 GB
32 GB
Memory type
HBM2
HBM2
Memory bandwidth
900 GB/s
900 GB/s
TDP
250 W
300 W
MSRP (list)
$8,000
$10,000
Buy
—
—
MSRP is the launch list price, not live retail. Buy links are Amazon affiliate links — the price after you click is the live price.