Access Before Ownership
Remote machine access, short-term loaners, cloud or lab credits, profiler access, and engineering contacts. This stage validates fixtures and reduces the risk of purchasing the wrong configuration.
C-Kernel-Engine needs more than benchmark claims. It needs repeatable measurements across SIMD generations, memory topologies, NUMA domains, and real network links. Support expands the laboratory while keeping the code, methods, limitations, and publishable results open.
The target is a ceiling for the complete multi-year program, not a requirement to spend everything immediately. Each CAD $5,000 milestone releases a more capable evidence lane only after the previous stage produces inspectable results.
Current funding must be reported manually. Until GitHub Sponsors and in-kind contributions are reconciled into a public ledger, these markers describe capability gates rather than claiming money raised.
This is money personally spent on hardware now used for CKE development and evidence generation. It is separate from sponsor funding, donated or loaned equipment, and the future CAD $30,000 laboratory ceiling.
| Hardware | Personal cost | Research role | Funding status |
|---|---|---|---|
| Lenovo ThinkPad T14 Gen 3 (used), Core i7-1260P, 24 GB DDR4-3200 (8 GB soldered + 16 GB SODIMM) | CAD $600 | Portable Ubuntu development and documentation workstation, AVX2 portability node, and interactive control node for server experiments | Personally funded |
| Lenovo P3 Tiny, Core i7-14700T (20 physical / 28 logical CPUs, AVX2/FMA/AVX-VNNI), 32 GB DDR5 | CAD $1,050 | Dedicated minimal Ubuntu Server node for repeatable CKE execution, profiling, and unattended experiments | Personally funded |
WD_BLACK SN850X 2 TB NVMe expansion for the P3 (CAD $499.99 + $35.00 PST + $25.00 GST), mounted at /data | CAD $559.99 | Artifact tier: model catalog, converted BUMP weights, isolated worktrees, X-Ray tensors, and VTune/Advisor/perf captures | Personally funded |
| Memory Express order WEB-1200399 (2026-08-11): Ryzen 9 9950X3D bundle — ASUS TUF X870E-PLUS WIFI7, G.SKILL Flare X5 64 GB DDR5-6000 (2×32 GB), Cooler Master 360 Elite AIO (incl. GST, BC PST and shipping) | CAD $2,470.87 | AMD evidence node core: native Zen 5 AVX-512 versus P3 AVX2/VNNI, cross-vendor numerics, distributed experiments — installed, verified on-node 2026-08-20, commissioning | Personally funded |
| Newegg order 574780872 (2026-08-11): MSI MPG A850GS 850 W PSU · Samsung 990 PRO 4 TB w/ heatsink · Corsair Vengeance 96 GB (2×48 GB) DDR5-6000 · ASUS TUF GT502 Horizon case · WD Red Plus 12 TB HDD · combo discount −$770.00 (incl. GST and PST) | CAD $4,177.52 | Ryzen node chassis, power, memory and storage tiers — installed, verified on-node 2026-08-20 (990 PRO 4 TB + WD Red 12 TB enumerated); the 96 GB Corsair kit is not in the current population; WD Red HDD for bulk artifact and corpus retention | Personally funded |
| Newegg order 574780892 (2026-08-11): Crucial 32 GB DDR5-5600 SODIMM (CT32G56C46S5) for the P3 (incl. ~12% GST+PST and $9.99 shipping) | CAD $716.78 | P3 second SODIMM — installed 2026-08-15, 64 GB total verified on-node; larger resident model lanes on the server node | Personally funded |
Ledger scope: this total includes only the six dated purchases above. The 2014 ThinkPad W530 remains available as a legacy node, but its original cost is not currently documented — it is marked pending verification and excluded from the total rather than estimated. The ledger also excludes electricity, personal labor, assembly, and future Xeon nodes. Future entries will separate personally funded, sponsor funded, donated, loaned, and discounted hardware.
Price-inflation note: the 2026-08-11 orders were paid during the AI-demand memory and NAND price spike — DDR5 kits and high-capacity NVMe in these orders cost multiples of their historical-normal prices (for example the Crucial 32 GB SODIMM at \$629.99 before tax and the 64–96 GB DDR5 kits above). The ledger records what was actually paid, not what the parts would have cost in a normal market.
The fleet is deliberately heterogeneous. Matched identical systems help controlled distributed scaling, but a mix of Intel generations, an AMD node, and one legacy machine exposes the ISA, scheduler, numerical, and provider-selection assumptions that identical hardware would silently hide. Each node below carries a status label: legacy/available, active portable, active server, ordered/assembly pending, or future target.
| Node | Status | CPU & ISA | Memory | Storage | OS role | Profiler | Research purpose |
|---|---|---|---|---|---|---|---|
| ThinkPad W530 (2014) | Legacy / available | Pending verification — old-x86, scalar/SSE/older AVX where supported | Pending verification | Pending verification | Legacy test node, remote orchestration | perf where supported | Old-x86 compatibility, scalar/SSE/older-AVX regression, failure and portability testing |
| ThinkPad T14 Gen 3 | Active portable | Core i7-1260P · AVX2 | 24 GB DDR4-3200 (8 GB soldered + 16 GB SODIMM) | Not an evidence tier | Ubuntu/AwesomeWM development and documentation workstation | perf (interactive) | Everyday CKE development, AVX2 portability, interactive control node for server experiments |
| Lenovo P3 Tiny | Active server | Core i7-14700T · 20 physical / 28 logical · AVX2/FMA/AVX-VNNI · no AVX-512 | 64 GB DDR5 (2×32 GB — second SODIMM Crucial CT32G56C46S5 installed 2026-08-15; verified on-node: 64,413,800 kB MemTotal) | 512 GB OEM system NVMe + 2 TB SN850X at /data (PCIe 4.0; 7,300 MB/s advertised) |
Minimal Ubuntu Server, unattended experiments | Intel VTune + Advisor (existing Intel evidence; Advisor is deprecated) + perf | Primary repeatable execution, profiling, and cross-family benchmark node |
Ryzen 9 9950X3D node (cke-amd-zen5-01) |
Active server — commissioning | Ryzen 9 9950X3D · Zen 5 · 16C/32T · full AVX-512 + BF16 + VNNI verified on-node 2026-08-20 · 128 MB L3 (96 MB 3D V-Cache) | 64 GB DDR5 (2×32 GB bundle kit) — verified 62,081,516 kB MemTotal; the 96 GB kit is not in the current population | Samsung 990 PRO 4 TB NVMe (3.6 TiB usable) + WD Red Plus 12 TB HDD — both verified on-node | Ubuntu 26.04 LTS · SSH evidence node (LAN + Tailscale) · 2.5GbE linked at 1 Gb/s | perf + AMD uProf (VTune hardware event sampling requires a genuine Intel CPU — not promised on AMD) | Native AVX-512 versus AVX2/VNNI, NVMe streaming and Memory Tetris research, cross-vendor numerics, distributed experiments with the P3 |
| Xeon 636 → 650/670 class | Future target | Xeon 630/650/670-class · AMX BF16 · AVX-512 | ECC RDIMM, four channels → eight with 650/670-class CPUs | NVMe per stage bill of materials | Sponsorship-roadmap lab node | VTune / Advisor / perf | AMX BF16 kernels, memory-channel scaling, matched-node distributed research |
The P3 keeps its original 512 GB Samsung OEM drive as the system disk and adds the SN850X as a dedicated artifact tier. The 7,300 MB/s figure is Western Digital's advertised sequential rating, not a CKE measurement; no dated fio artifact is published for either drive yet. The third tier shown is the Samsung 990 PRO 4 TB installed in the Ryzen node (verified on-node 2026-08-20: 3.6 TiB usable) — a Gen4 drive at 7,450 MB/s advertised; the earlier Gen5 9100 PRO quote was not purchased.
Roughly 1.8 TiB of usable space (ext4, noatime, mounted read/write at /data) changes how much evidence the lab can keep, not how fast the machine computes:
Scope limit: capacity improves experimental breadth and artifact retention. It does not remove RAM constraints, and no storage tier makes an SSD equivalent to DRAM.
Status: installed and currently being tested. The node (cke-amd-zen5-01, reachable over LAN and Tailscale) passed its verification packet on 2026-08-20: lscpu confirms the Ryzen 9 9950X3D at 16 cores / 32 threads on one NUMA node with the full Zen 5 AVX-512 flag set (F/BW/CD/DQ/VL, VNNI, BF16, IFMA, VBMI/VBMI2, BITALG, VPOPCNTDQ, VP2INTERSECT), AVX-VNNI and FMA; cache topology L1d 48 KB + L1i 32 KB per core, L2 1 MB per core, L3 128 MB total (96 MB 3D V-Cache CCD + 32 MB); 64 GB DDR5 installed (62,081,516 kB MemTotal); Samsung 990 PRO 4 TB (3.6 TiB usable) and WD Red Plus 12 TB (10.9 TiB) both enumerated; onboard Realtek 2.5GbE currently linked at 1 Gb/s. The two dated 2026-08-11 orders below remain the cost record. Baseline stability and first CKE experiments are the commissioning step now in progress; no performance claim is made until measured and published.
| Order | Contents | Paid |
|---|---|---|
| Memory Express WEB-1200399 (receipt, paid via Adyen) | AMD Ryzen 9 9950X3D (16C/32T, up to 5.7 GHz) · ASUS TUF Gaming X870E-PLUS WIFI7 · G.SKILL Flare X5 64 GB DDR5-6000 CL36 (2×32 GB) · Cooler Master MasterLiquid 360 Elite | CAD $2,470.87 |
| Newegg 574780872 + 574780892 (order confirmation) | MSI MPG A850GS 850 W PSU · Samsung 990 PRO 4 TB w/ heatsink (PCIe 4.0 x4) · Corsair Vengeance 96 GB (2×48 GB) DDR5-6000 · ASUS TUF GT502 Horizon case · WD Red Plus 12 TB HDD · Crucial 32 GB DDR5-5600 SODIMM (for the P3) | CAD $4,894.30 |
Paid figures are the order totals including GST, PST and shipping; the combo discount (−\$770.00) is already applied in the Newegg total. These orders replaced the 2026-08-09 quotes, which are kept here for the record: the Memory Express bundle quoted CAD \$2,174.98 (actually paid \$2,470.87 including tax and shipping) and the Newegg cart estimated CAD \$1,319.94 before tax for a smaller basket — the final order grew to include the 96 GB DDR5 kit, the 12 TB HDD, and the P3 SODIMM. Samsung's 7,450 MB/s sequential figure for the 990 PRO 4 TB is a vendor rating until measured on the final motherboard; the earlier 9100 PRO Gen5 quote (14,800/13,400 MB/s advertised) was not purchased. CPU claims follow the AMD product page: Zen 5, 16 cores / 32 threads, AVX-512 and AVX2, 128 MB L3, PCIe 5.0, dual-channel DDR5, 170 W. Memory pricing in these orders reflects the 2026 AI-demand spike — see the ledger's inflation note above.
Research role: native Zen 5 AVX-512 versus P3 AVX2/AVX-VNNI comparisons; the current memory population is the 64 GB bundle kit (the 96 GB Corsair kit remains on the shelf for a later re-population); NVMe streaming and Memory Tetris experiments on the 990 PRO tier; cross-vendor numerical repeatability and provider selection; larger Gemma, GLM, Kimi, Qwen3.6, Whisper, and multimodal artifact-backed tests where memory permits; and distributed CKE experiments with the P3 as the Intel node and Ryzen as the AMD node — no scaling claim until measured. On AMD, Linux perf and AMD uProf (hardware counters, event-based profiling, IBS) are the primary PMU tools; Intel documents VTune hardware event-based sampling as requiring a genuine Intel processor, so no VTune PMU analysis is promised on Ryzen, and the deprecated Advisor remains an Intel-only evidence tool. This node is an intermediate founder-funded AVX-512 and storage lane — it does not replace the Xeon AMX/BF16/memory-channel sponsorship roadmap below.
Founder funding establishes matched Xeon 636 nodes for native AMX BF16, AVX-512/VNNI, four-channel memory, and direct-network experiments. Sponsor support then unlocks the Xeon 600 platform's full eight-channel path through compatible 650/670-class processor upgrades and four additional matched RDIMMs per node. The resulting four-channel-versus-eight-channel comparison becomes a published experiment rather than an assumed benefit.
Remote machine access, short-term loaners, cloud or lab credits, profiler access, and engineering contacts. This stage validates fixtures and reduces the risk of purchasing the wrong configuration.
A balanced workstation with ECC memory across the four channels exposed by the 630-series CPU, NVMe storage, sustained cooling, power measurement, and room for a high-speed NIC. It unlocks native AMX BF16 and controlled memory work.
A second matched node plus two 100GbE-capable NICs, a direct DAC link, storage, metering, and contingencies. A switch is deferred until a third node creates a measured need.
The W890 workstation platform has eight DIMM slots, but ASUS specifies that a Xeon 630-series processor enables only channels A, C, E, and G. The most valuable hardware sponsorship is therefore two compatible Xeon 650/670-class processors plus four additional matched RDIMMs per node. CKE will publish the controlled four-channel-versus-eight-channel result across memory bandwidth, prefill, decode, AMX BF16 kernels, power, NUMA behavior, and distributed workloads. Final CPU selection will also score cache per core, core count, topology, price, and availability; channel count alone will not decide the configuration.
Budget discipline: these are planning ranges, not a final bill of materials. CPU availability, ECC-memory pricing, tax, cooling, chassis, PSU connectors, NIC provenance, and warranty must be quoted before purchase. Funds are released by stage; a second node is not purchased until the first-node evidence pipeline is working.
| Experiment lane | Measurements | Required evidence |
|---|---|---|
| AMX BF16 kernels | Tile utilization, conversion boundaries, accumulation semantics, throughput, and parity against PyTorch | Commands, shapes, ISA detection, assembly/profiler artifacts, numerical tolerances |
| AVX-512 and VNNI | FP32 and quantized GEMM/GEMV, attention, reductions, fusion, cache behavior | AVX2 baseline, compiler flags, warmups, repeats, end-to-end impact |
| Memory channels | One through fully populated channel configurations where practical; bandwidth, latency, and model throughput | DIMM topology, frequency, NUMA map, measured STREAM-like baseline and real CKE workload |
| NUMA placement | First-touch, interleave, binding, page size, thread affinity, local versus remote access | numactl topology, allocation policy, counters, RSS, latency and throughput |
| Single-node inference | Prefill, decode, time to first token, RSS, power, thermals, numerical parity | Model hash, quantization, context, batch, threads, commit, reference runtime |
| Bounded training | Forward/backward parity, optimizer state, memory budget, BF16/FP32 behavior and step time | Explicit supported circuit and dataset; no extrapolation to frontier-scale training |
| Distributed execution | One-node versus two-node scaling, tensor/pipeline partitioning, communication overlap, replicated versus sharded state | Topology, payload sizes, synchronization, transport, efficiency and failure modes |
| Network crossover | 10/25/100GbE or available links; sockets/MPI and RDMA where supported | Wire rate, CPU cost, message-size sweep, latency, bandwidth and end-to-end model effect |
| Power and TCO | Idle, kernel, model, and distributed wall power; energy per workload | Meter model, measurement interval, local electricity assumption and dated BOM |
| Portability | Core i7, Xeon generations, ARM/NEON and future AMD/Arm server access | Same fixtures and canonical output boundaries; architecture-specific limitations retained |
Kernel and runtime patches, parity gates, benchmark commands, machine-readable result records, and versioned hardware metadata.
VTune, Advisor, perf, assembly, memory, network, power, and end-to-end reports with failures and negative results retained.
Source-linked documentation, ShivasNotes articles, diagrams, and videos that explain what changed, why it matters, and where the result does not generalize.
Use GitHub Sponsors for direct financial support. Use ShivasNotes Work With Us for a paid private investigation with bounded deliverables. Hardware vendors and laboratories can contact Anthony Shivakumar about loaners, remote access, parts, or discounts.
No fleet result — from the W530 to the future Xeon nodes — is published without the same minimum evidence packet. A number without this packet is a lab note, not a CKE result.
lscpu output and detected ISAnvme list / device SMART stateSponsorship does not purchase a favorable conclusion. Every published result identifies hardware, firmware where available, detected ISA, memory topology, operating system, compiler, CKE commit, model hash, quantization, workload, thread policy, commands, warmups, repetitions, reference backend, date, and known limitations.
Financial support, loaners, discounts, credits, and vendor engineering assistance are disclosed beside affected results. Sponsors may review factual hardware details before publication but cannot suppress reproducible negative findings or rewrite conclusions. Security-sensitive credentials and legitimately confidential information are excluded from public artifacts.
Send the hardware model, access constraints, support type, available dates, and the question you want tested to anthony.shivakumar@antshiv.com. For paid engineering work with private inputs and bounded deliverables, use the Work With Us page.