Model Sizing Lab

Six accelerators do not form one large memory pool — and neither do six CPU sockets. Place a model, choose its active contexts, and see where weights, KV cache and recurrent state consume capacity on GPUs or CPUs.

Estimator, not a benchmark or service guarantee Theoretical bandwidth ceilings only
SIX INDEPENDENT MEMORY BANKS BANK 1BANK 2BANK 3BANK 4BANK 5BANK 6

1. Choose a model

Hugging Face lookup reads repository metadata only. Select the exact GGUF or safetensors files; no weight download occurs.

Preset patterns are illustrations, not verified checkpoint configurations. Inspect and correct every field.

Loading CKE's serving coverage inventory…

Disk bytes are a first approximation to loaded weight bytes. Conversion, repacking, scales and padding can change this.

2. Declare the cache

Enter distinct KV storage owners, not just the number of attention layers. A layer reusing another layer's cache does not own another copy.

Eviction is a serving implementation choice, not guaranteed by sliding attention. CKE currently plans full-context Gemma 4 KV even where its kernel reads only the window.

3. Workload & hardware

CPU capacity is system RAM usable by the process; GPU capacity is device memory. All preset bandwidths are theoretical peaks — derate with the efficiency input. On a CPU, one socket's weights are shared by all its requests; sockets still do not pool.

Fit on one replica

-At chosen concurrency
-Estimated used memory
-Headroom / deficit
-Selected weight bytes
-KV + state / request
-Approx. max active requests
WeightsKVRecurrent stateReserveFree

Decode bandwidth ceiling

Optimistic thought experiment

-Upper bound, tokens/s per active request
-Upper bound, tokens/s per replica
-Upper bound across selected replicas

batch bytes ≈ weights + active requests × (live KV read + recurrent state read)
aggregate token/s ≤ active requests × effective bandwidth ÷ batch bytes

Assumes one shared weight read per batch step, perfect scheduling, and all modeled state read once. It excludes compute, dequantization, interconnect, network, prefill, queueing and many kernel costs. A number here is not expected throughput. It is not valid when the request does not fit. On CPUs this bound is usually dominated by DRAM bandwidth — see Serving & Batching for what concurrency CKE actually executes today.

Replica placement

Mark GPUs or CPU nodes serving replicas of this model; reserve the others for OCR, ASR or different models. Each replica carries its own weights and request state.

Inspect the assumptions

What this page cannot infer

File bytes do not identify attention geometry, loaded weight precision, engine workspace or real utilization. Model-family presets are patterns only. Hugging Face metadata may be missing, gated, incomplete or ambiguous. Preprocessing and vision/audio towers need their own model memory and latency budgets.

What the CKE circuit picker knows

The picker reads CKE's generated serving coverage inventory. A circuit can declare a topology pattern (for example a 3:1 recurrent-to-full-attention hybrid block structure); it never declares numeric dimensions. Where a ratio pattern exists, a total layer count from Hugging Face metadata is split into full-attention owners and recurrent layers automatically — labeled as an estimate, not a certification.

How to move from bound to service envelope

Benchmark the exact checkpoint, serving backend, context distribution and mixed workload. Measure p95 first-token time, per-user token rate, queue time, and aggregate throughput. Compare observed allocated memory with this worksheet and with CKE's memory plan. See the CKE roofline discussion for why prefill and decode differ.

Hardware references: NVIDIA RTX PRO 6000 Blackwell Server Edition. Hugging Face metadata source: Hub API. CKE circuit inventory: v8 Serving Coverage. This page runs in your browser and does not send your scenario to CKE.

Image
100% | |
Scroll to zoom | Drag to pan | W/H to fit | 0 to reset | ESC to close