v8 Kernel Architecture
The v8 engine keeps every kernel behind a provider map: a JSON document in
version/v8/kernel_maps/ that names the operation, the numerical contract, the port layout, the C
call ABI, and the selection metadata. The registry currently holds 350 provider maps across 124
operations. Circuits choose what runs — a logical operation with a numerical contract; the maps
decide which physical kernel satisfies it on this host.
Registry at a Glance
350 provider maps, 124 logical operations. Attention dominates because every dtype × phase ×
layout combination is a separate physical provider with its own contract; the long tail is singleton
operations such as final_logit_scale or attn_gate_softplus_mul that exist because one
model family needed exactly that arithmetic.
The schema also defines
diagnostic and deprecated; no current map uses them.
Explicit candidate providers resolve only behind explicit opt-in routes (for example
production_when_prepared weight layouts), and equal-priority production ambiguity fails
closed instead of guessing. The 260 unmarked maps are implicit legacy providers: they
remain eligible as compatibility fallbacks and rank below production — migration debt, not
unreachable providers.
moe_softmax_topk_router_llama_f32,
moe_softmax_topk_router_pytorch_bf16,
moe_swiglu_expert_forward_q4k_q5k_bucketed,
moe_swiglu_expert_forward_nvfp4,
moe_swiglu_shared_forward_q8_0_gated,
moe_swiglu_shared_forward_nvfp4,
geglu_forward_ggml_native,
ple_gate_conv_inject_llama_fp16,
farskip_swiglu_shared_combine_bf16 — each waits for measured evidence before promotion.
(moe_swiglu_expert_forward_q4k_q5k was promoted to production at priority 200; the bucketed
variant remains candidate at 210.)
version/v8/kernel_maps/KERNEL_REGISTRY.json is regenerated by ck_run_v8.py
(never hand-edited). The migration ratchet is
version/v8/contracts/kernel_interface_migration_baseline.json, enforced by
audit_kernel_map_interfaces_v8.py --check in CI. Current audit census: hardened 84,
interface_abi 84, map_abi 223, selection_managed 90, contract_pending 49, legacy 211, legacy_interface 57,
legacy_if 59, op_if 28 (of 350 maps). See
Kernel Maps and Provider Selection for the map schema itself.
Snapshot: version/v8/kernel_maps/KERNEL_REGISTRY.json (350 maps) +
version/v8/scripts/audit_kernel_map_interfaces_v8.py at commit 9284499dc (2026-09).
These counts are regenerated by hand against audit output until the docs-build generator exists.
How a Provider Is Selected
Selection is a filter chain, not a score. Compatibility gates run first; priority only ranks providers that already implement the same logical and numerical contract. Priority can never promote a candidate over a production provider, pick across equivalence groups, or override dtype, layout, phase, shape, ISA, or ABI compatibility.
The Q4_K × Q8_K AVX-512 VNNI x16 provider follows the same discipline: prefill routes are
candidate behind CK_V8_FORCE_BATCHED_PREFILL=1 until 1K logit parity is certified;
only the decode (m=1) prepared route is production_when_prepared.
Attention — 42 providers
One logical operation, many physical contracts. Prefill and decode are separate providers with separate phases; flash online-softmax and exact full-score attention sit in different equivalence groups because their reduction order differs. The base math (QKᵰ/√d, causal mask, softmax, online-softmax tiles) is derived in Deep Dive Concepts and Flash Attention Analysis; this section covers only the variants the registry distinguishes.
attention_forward_causal_head_major_gqa_flash_strided production
(priority 200, group attention.fp32_online_causal_strided.v1) and the decode flash family
(_f16kv, _f16cache, _bf16cache_pytorch_contract). Head-major layout,
FP32 online-softmax accumulation, runtime ISA dispatch (AVX-512 / AVX2 / AVX / reference).
attention_forward_*_sliding[_gemma4] (op attention_sliding, 9 maps) — same flash
structure with a banded mask; used by Gemma3/4, Cohere2 (3:1 sliding:full), and Laguna's sliding MoE
layers. Mask math: concepts — sliding-window attention.
Qwen3.5/3.8 apply an elementwise sigmoid gate (
attn_gate_sigmoid_mul_forward):attention_forward_{causal,decode}_head_major_shared_kv[_sliding]_gemma4 — one KV stream shared
across head groups, selected from circuit data rather than a family branch.
The softplus gate (attn_gate_softplus_mul_forward, hybrid_attention_kernels.c) is a
per-head scalar — one gate value per head, with a linear branch above 20 for numerical safety. The
sigmoid gate is per-element. They share an op shape but not a numerical contract, so they are separate
providers, never an equivalence group. Laguna's softplus formula is observed from the model, not yet
oracle-validated.
Multi-head latent attention (MLA)
Kimi and Instella cache a compressed latent KV vector, not per-head K/V. Decode and prefill insert
explicit cache ops (mla_kv_cache_store / mla_kv_cache_batch_store), and the attention
provider decompresses on the fly:
Providers: deepseek_mla_attention_f32 / _decode_f32,
deepseek_mla_kv_decompress_{f32,bf16},
deepseek_mla_partial_rope_concat_{f32,packed_f32,packed_bf16_storage} — all in
src/kernels/deepseek_kernels.c. Cache rows are zero-padded from head_dim to cache_stride. Full
contract walkthrough: v8 MLA / Kimi decode cache.
RoPE — 18 providers + YaRN + M-RoPE
Every rotary layout is a distinct contract because the channel pairing is a checkpoint semantic, not a tuning knob. The base rotation math and the split-half vs pairwise layouts are derived in concepts — RoPE; the registry adds three more axes:
rotary_dim < head_dim: channels
[rotary_dim, head_dim) pass through unrotated (GLM-4, GPT-NeoX-style partial). Implemented by
the *_with_rotary_dim providers.rope_forward_qk_split_direct_f32 recomputes
frequencies per layer from IR metadata (rope_param_mode: per_layer_direct) instead of one
global cache.yarn_rope_cache_* (3 providers, op
yarn_rope_init): ramp-blended inverse frequencies,
inv_freq = inv_interp·ramp + inv_extrap·(1−ramp), with correction-dim and mscale from GGUF
metadata. Laguna mixes this with full-rotary sliding layers in one model.mrope_qk_* sectioned per-axis
positions with YaRN correction per pair; multimodal_mrope_positions_2d builds the position
tables. Qwen3-VL, Qwen3.8 text, Gemma4 Vision.
ULP-level note: the llama.cpp-exact provider (rope_precompute_cache_llama_cpu) builds the angle
table by iterative FP32 multiplication (theta *= theta_scale) instead of per-pair
powf, and the M-RoPE contract resolves the system libm cosf/sinf/powf via dlopen to
avoid libimf divergence on ICX hosts. Bit-exactness against an oracle is a property of these details, which is
why RoPE has 18 providers and not 3.
Norms — 28 providers
v_norm for Gemma4.recurrent_norm_gate_* — RMSNorm fused
with the SiLU output gate inside DeltaNet blocks; mamba2_rmsnorm_gate_f32 for Nemotron-H.Quantized GEMM / GEMV — 38 providers
Weight-only quantized matmul with BF16/FP32 or Q8_K activations. Format details live in Quant Fundamentals and GEMM Memory Layout; the registry-level facts that matter for selection:
| Contract | Providers | Notes |
|---|---|---|
| Q4_K × Q8_K | gemm_nt_q4_k_q8_k, gemv_q4_k[_q8_k] |
AVX-512 VNNI x16 packed routes: decode (m=1) prepared is production_when_prepared; prefill routes are candidate behind CK_V8_FORCE_BATCHED_PREFILL=1 pending 1K parity. |
| Q6_K × Q8_K | gemm_nt_q6_k_q8_k, gemv_q6_k[_q8_k] |
Exact compact Q6 is production. The prepared expanded layout (dequant to weight = d·scale·(q6−32) with ql/qh pre-merged at load) is candidate since PR #405 — measured 0.923× on Ryzen. |
| Q8_0 / Q5_x / Q4_0/1 | gemm_nt_q8_0*, gemm_nt_q5_*, gemv_* |
Plus fused GEMV with online quant + bias (gemv_fused_q5_0_bias, gemv_fused_q8_0_bias). |
| BF16 / F16 | gemm_nt_bf16[_native|_amx|_pytorch_onednn_brgemm], gemm_nt_f16[_clipped] |
AMX and oneDNN BRGEMM routes for Xeon; BF16 storage contracts for the PyTorch parity lanes. The Cohere Compass BF16 frontend added gemv_bf16_parallel_dispatch (thread-pool row-partitioned decode GEMV) and a parallel gemm_nt_bf16_bf16_storage variant with exact BF16 round-trip storage. |
| Exact FP32 | gemm_nt_fp32_exact, gemm_nt_f32_llama_production |
The exact parallel row providers (ck_parallel_prefill_v8.c) that the Qwen3.6/3.8 long-prefill lanes use when parity outranks speed. |
| Head-major projection | qkv_projection, attention_projection |
ck_qkv_project_head_major_quant / ck_attention_project_head_major_quant write head-major outputs directly — no layout conversion pass between projection and attention. |
MoE — routers and experts
Four model families route: Nemotron-H (group-limited top-k + ReLU2), Kimi and Instella (sigmoid group-limited
top-k + shared SwiGLU), Laguna (sigmoid + correction bias, top-8 + shared, mixed Q4/Q6 expert storage).
The selection rule, verified from src/kernels/topk_kernels.c:
nemotron_group_limited_topk_router_f32 and
group_limited_topk_router_sigmoid_f32 (same rule, sigmoid-scored).
moe_softmax_topk_router_llama_f32 (full softmax then top-k, floor 6.1e-5) is
candidate.moe_relu2_expert_*) for Nemotron-H;
SwiGLU routed + shared (moe_swiglu_expert_*, moe_swiglu_shared_*) for Kimi /
Instella / Laguna. Production quant contracts: q4k_q5k (priority 200), q4k_q4k and q4k_q6k (priority 205,
with *_parallel_workspace thread-pool variants from PR #422); the bucketed grouped-prefill
variant (priority 210) and q8_0_gated shared are candidate. FarSkip
shared-combine (farskip_swiglu_shared_combine_bf16) is
candidate. Kernel-by-kernel walk:
MoE Expert Kernels.Recurrent / SSM — DeltaNet and Mamba2
Full derivations live in concepts — recurrent attention & Mamba2 and the Gated DeltaNet deep dive. The registry-level view:
..._parallel_forward), llama-AVX2 prefill dispatch, and PyTorch-grouped BF16
storage variants — src/kernels/deltanet_kernels.c. Qwen3.5, Qwen3.6, Qwen3.8.src/kernels/mamba2_kernels.c.
Nemotron-H. State is shaped [heads, head_dim, state_dim], not a square DeltaNet matrix.recurrent_gate, recurrent_silu,
recurrent_qk_l2_norm, recurrent_split_(conv_)qkv,
recurrent_conv_state_update, ssm_conv1d_* — the plumbing that keeps conv state and
split layouts explicit in the IR instead of hidden inside a monolith kernel.KV Cache — persistent state, valid rows vs capacity
Cache tensors are persistent state ports, not ephemeral outputs. The layout is head-major
[kv_head, token, aligned_head_dim]; capacity (max_seq_len) and valid-token count are
declared separately, and kernels must never read beyond valid rows — scheduling may round to a physical
extent, the kernel may not.
kv_cache_store (f32),
_f16, _bf16, batch f16/bf16, and kv_cache_store_shared_q for Gemma4
shared KV. Single-token stores guard pos < max_seq_len; batch stores guard
start_pos + num_tokens ≤ max_seq_len.kv_cache_repack_head_major_inplace clamps
tokens = min(tokens, cache_capacity) and memmoves head blocks high→low when capacity grows —
the memory planner can resize the arena without invalidating live state.deepseek_mla_kv_cache_{store,batch_store}) keep head-major rows zero-padded to
cache_stride.Fused providers — earned, not assumed
src/kernels/fused/ holds 11 fusion sources. Fusion is a measured decision, not a default: the
policy (see Kernel Reference — composed vs fused) keeps operations composed in the
IR unless profiling proves the fused epilogue is worth the maintenance. Registered examples:
mega_fused_attention_decode_q5_0 (9 ops fused),
mega_fused_attention_prefill[_q8_0] (RMSNorm → QKV → RoPE → flash → out-proj + residual),
fused_mlp_block (OutProj → residual → RMSNorm → MLP → residual),
fused_rmsnorm_qkv_prefill_head_major_quant, and
rmsnorm_q8_k_fused (norm + activation quantization).
Logits and footer ops
final_logit_scale_f32 (logit_kernels.c):
in-place logits[i] *= scale. Cohere2 declares logit_scale in its GGUF metadata; the
circuit carries it and lowering appends this footer op — no family branch in codegen.gemma4_final_logit_softcap_forward:
logits = tanh(logits / cap) · cap — Gemma-style softcapping as an explicit op.assistant_layer_scale_forward and
logits_copy_to_position: Gemma4 assistant-layer scaling and decode-time
logits placement — small ops that exist because correctness lives in the details.Audio, vision, and training
audio_preemphasis_f32,
audio_stft_power_centered_window_f32, audio_log_mel_time_major_f32,
audio_feature_normalize_per_feature_f32, audio_conv2d_whc_grouped_f32,
audio_glu_split_channel_major_f32, audio_relative_shift_f32 — contract-tested but
not selection-managed. FP32 throughout; walkthrough in
concepts — audio frontend,
Audio Kernels Deep Dive, and
Cohere Kernel Story — Transcribe.im2patch, BF16 patch projection (oneDNN conv3d
storage, plus patch_projection_image_bf16_native_storage — thread-pooled, temporal-2 weight
pair, merge-aligned grid, BF16 native storage for Cohere Compass), spatial merge (2×2 / tiled /
average-pool), 2D position ids, multimodal prefix insert.
Deep dive: v8 Vision Encoder Architecture.adamw_update and softmax cross-entropy loss. The v7 training lane:
v7 Inference + Training Runbook.Composite Circuits — components + stitch
Multimodal families are built by composing certified components rather than writing a monolithic new
circuit. Circuits compose three ways: extends (a child deep-merges over a parent circuit, e.g.
cohere_compass_text extends cohere2), components (a circuit embeds named
sub-circuits, e.g. cohere_compass embeds decoder → cohere_compass_text
and vision_encoder → cohere_compass_vision), and stitch edges
(declared tensor handoffs between components, each carrying its own
required_contract.providers). The compiler resolves extends/components
recursively and fails loudly on cycles or unresolvable references.
cohere_compass_text inherits cohere2's layer plan and overrides only what the
Compass decoder changes (M-RoPE text positions, the visual-prefix input). No copy of the parent graph, no
forked maintenance.cohere_compass.json names two sub-circuits with
explicit exports/imports: the vision encoder publishes
vision_embeddings, the decoder consumes visual_prefix. Each component keeps its own
provider bindings and certification status.vision_embeddings_to_decoder_prefix (op multimodal_prefix_stitch) declares the
handoff contract — token-major FP32 vision storage, M-RoPE 2D positions, mixed visual/text prefill
— and binds three providers: multimodal_prefix_insert_f32,
multimodal_mrope_positions_2d, mrope_qk_imrope_positions.python3 version/v8/scripts/report_model_novelty_v8.py --circuit cohere_compass reports the
composite attribution above (37 ops / 24 providers; per-component rows; stitch edge providers); the report
has been composite-aware since PR #462 (fix commit cbc8296a2).
Circuit sources: version/v8/circuits/cohere_compass.json (components + stitch),
cohere_compass_text.json (extends cohere2), cohere_compass_vision.json. The same
shape powers qwen36vl.json; kimi_vl.json is a flat circuit for contrast. Bring-up
story: Cohere Kernel Story — North Micro Vision / Compass.