← AI Hashrate 中文

Can V100 16GB PCIe run Hy3?

NVIDIA · 16 GB HBM2 · 900 GB/s bandwidth · Tencent · 295.0B (active 21.0B) · MoE · model ctx up to 256K

❌ No — Hy3 Q4 needs 164.5 GB, card only has 16 GB
VRAM: 164.5 GB needed of 16 GB
FF · No FitComposite score: VRAM headroom × 0.5 + tok/s speed × 0.5. Q4 @ 8K context.
CUDA
Your GPU won't fit Hy3, but runs these similar models: Qwen3-8B, gemma-4-26B-A4B-itor☁️ Rent on RunPod from ~$0.50/hr
Estimates from memory-bandwidth formula; rows marked "measured" override. Methodology.

Fit & speed by quant and context

QuantContextVRAM neededFits?tok/s (decode)
Q44K164.5 GBNo
Q48K default165.75 GBNo
Q432K173.25 GBNo
FP164K592.25 GBNo
FP168K default593.5 GBNo
FP1632K601.0 GBNo

Fits = weights + KV(ctx) + 1 GB overhead ≤ 95% of VRAM. Measured anchors are context-agnostic; the fit verdict is recomputed per context. A missing tok/s means the model is far beyond this card (offload-only territory).

VRAM breakdown at 8K context

QuantWeightsKV cacheOverheadTotal neededV100 16GB PCIe VRAM
Q4162.25 GB2.5 GB1 GB165.75 GB16 GB
FP16590.0 GB2.5 GB1 GB593.5 GB16 GB

Other GPUs that run Hy3

All GPUs for Hy3 →

Other models for the V100 16GB PCIe

All models on the V100 16GB PCIe →