← AI Hashrate 中文

Can V100 32GB SXM2 run Inkling?

NVIDIA · 32 GB HBM2 · 900 GB/s bandwidth · Thinking Machines · 975.0B (active 41.0B) · MoE · model ctx up to 1024K

❌ No — Inkling Q4 needs 538.28 GB, card only has 32 GB
VRAM: 538.28 GB needed of 32 GB
FF · No FitComposite score: VRAM headroom × 0.5 + tok/s speed × 0.5. Q4 @ 8K context.
CUDA
Your GPU won't fit Inkling, but runs these similar models: Qwen3-8B, gemma-4-26B-A4B-itor☁️ Rent on RunPod from ~$0.50/hr
Estimates from memory-bandwidth formula; rows marked "measured" override. Methodology.

Fit & speed by quant and context

QuantContextVRAM neededFits?tok/s (decode)
Q44K538.28 GBNo
Q48K default539.31 GBNo
Q432K545.5 GBNo
FP164K1952.03 GBNo
FP168K default1953.06 GBNo
FP1632K1959.25 GBNo

Fits = weights + KV(ctx) + 1 GB overhead ≤ 95% of VRAM. Measured anchors are context-agnostic; the fit verdict is recomputed per context. A missing tok/s means the model is far beyond this card (offload-only territory).

VRAM breakdown at 8K context

QuantWeightsKV cacheOverheadTotal neededV100 32GB SXM2 VRAM
Q4536.25 GB2.06 GB1 GB539.31 GB32 GB
FP161950.0 GB2.06 GB1 GB1953.06 GB32 GB

Other GPUs that run Inkling

No catalog GPU fully fits Inkling at Q4/8K — see the model page for offload estimates. model page

All GPUs for Inkling →

Other models for the V100 32GB SXM2

All models on the V100 32GB SXM2 →