The RTX 5090 32GB has about 833% more memory bandwidth (1,792 vs 192 GB/s). Batch-1 decode tracks bandwidth closely, so a model that fits both cards should decode roughly 833% faster on the RTX 5090 32GB. The RTX 5090 32GB fits more of the catalog at Q4/8K — 39 vs 16 of 69 models (32 vs 8 GB VRAM). Largest model that fits the RTX 5090 32GB but not the Tesla P4 8GB: Qwen3.6-35B-A3B (36.0B). MSRP is $80 for the Tesla P4 8GB vs $1,999 for the RTX 5090 32GB. MSRP is a launch list price, not live retail — Amazon prices move; click through for the current price. In the Q4/8K head-to-head below (39 shared models), the RTX 5090 32GB posts the higher tok/s in 39 rows vs 0. Methodology.