ADR 0001: CPU kernel strategy¶
Status: Accepted. Scope: src/core/kernels/cpu/.
Context¶
Level III kernels for the forward pass. Targets AVX2 on x86 edge and NEON on Jetson. Parity with PyTorch is required, so fast-math is off. Kernels are allocation-free and run on arena buffers.
Decision¶
Explicit intrinsics for the kernels that carry the runtime:
matmul,silu,softplus,gate_silu, and fuseddiscretize_and_scan.expandloguse hand-written Cephes polynomials, about 1 ULP. The compiler folds a scalarexpinto a per-element libm call, so intrinsics are the only way to vectorize it. Constants are shared between the ISA files incephes.h.conv1d_causal,conv1d_stepandrmsnormstay simple scalar loops inside each backend. They are under 3% of a layer, too little to dominate the profile.Scan layout is state-major
[t][n][c]so the inner channel loop is contiguous.
Consequences¶
Portable. A scalar fallback compiles where SIMD is absent.
matmulis the dominant cost and the first optimization target.The
expapproximation is a bounded, deliberate deviation. Full parity is enforced by the ULP gate.