Model Sizing Lab
Six accelerators do not form one large memory pool — and neither do six CPU sockets. Place a model, choose its active contexts, and see where weights, KV cache and recurrent state consume capacity on GPUs or CPUs.
1. Choose a model
Hugging Face lookup reads repository metadata only. Select the exact GGUF or safetensors files; no weight download occurs.
Preset patterns are illustrations, not verified checkpoint configurations. Inspect and correct every field.
Loading CKE's serving coverage inventory…
Disk bytes are a first approximation to loaded weight bytes. Conversion, repacking, scales and padding can change this.
2. Declare the cache
Enter distinct KV storage owners, not just the number of attention layers. A layer reusing another layer's cache does not own another copy.
Eviction is a serving implementation choice, not guaranteed by sliding attention. CKE currently plans full-context Gemma 4 KV even where its kernel reads only the window.
3. Workload & hardware
CPU capacity is system RAM usable by the process; GPU capacity is device memory. All preset bandwidths are theoretical peaks — derate with the efficiency input. On a CPU, one socket's weights are shared by all its requests; sockets still do not pool.
Fit on one replica
Decode bandwidth ceiling
Optimistic thought experiment
batch bytes ≈ weights + active requests × (live KV read + recurrent state read)
aggregate token/s ≤ active requests × effective bandwidth ÷ batch bytes
Assumes one shared weight read per batch step, perfect scheduling, and all modeled state read once. It excludes compute, dequantization, interconnect, network, prefill, queueing and many kernel costs. A number here is not expected throughput. It is not valid when the request does not fit. On CPUs this bound is usually dominated by DRAM bandwidth — see Serving & Batching for what concurrency CKE actually executes today.
Replica placement
Mark GPUs or CPU nodes serving replicas of this model; reserve the others for OCR, ASR or different models. Each replica carries its own weights and request state.
Inspect the assumptions
What this page cannot infer
File bytes do not identify attention geometry, loaded weight precision, engine workspace or real utilization. Model-family presets are patterns only. Hugging Face metadata may be missing, gated, incomplete or ambiguous. Preprocessing and vision/audio towers need their own model memory and latency budgets.
What the CKE circuit picker knows
The picker reads CKE's generated serving coverage inventory. A circuit can declare a topology pattern (for example a 3:1 recurrent-to-full-attention hybrid block structure); it never declares numeric dimensions. Where a ratio pattern exists, a total layer count from Hugging Face metadata is split into full-attention owners and recurrent layers automatically — labeled as an estimate, not a certification.
How to move from bound to service envelope
Benchmark the exact checkpoint, serving backend, context distribution and mixed workload. Measure p95 first-token time, per-user token rate, queue time, and aggregate throughput. Compare observed allocated memory with this worksheet and with CKE's memory plan. See the CKE roofline discussion for why prefill and decode differ.
Hardware references: NVIDIA RTX PRO 6000 Blackwell Server Edition. Hugging Face metadata source: Hub API. CKE circuit inventory: v8 Serving Coverage. This page runs in your browser and does not send your scenario to CKE.